Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.
Ma et al., ACL 2026
Key facts
| What it measures | Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models. |
| Who should care | Anyone deciding which benchmark to trust before they trust its scores. |
| Use it when | You want to know whether a benchmark deserves attention at all. |
| Do not use it for | You want model rankings. It has none. |
| Score scale | Percent of checklist criteria met, per phase. |
| Grader method | Three independent NLP researchers, discrepancies resolved by consensus |
| First released | 2025 |
| Languages | English |
| Use cases | meta |
| Citation | Ma et al., ACL 2026 |
| Status | active |
Official site · Code · Paper · DOI
Independence assessment
Academic, four universities, no commercial sponsor identified. Audits benchmarks but does not track ownership, funding or steward conflicts, which is the gap this registry fills.
Scope
Five-phase, 46-criteria lifecycle framework applied to 53 medical LLM benchmarks (56 in the ACL version). Base instrument adopted by this registry.
Frequently asked questions
What does MedCheck measure?
Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.
Who should use MedCheck?
Anyone deciding which benchmark to trust before they trust its scores. You want to know whether a benchmark deserves attention at all.
What should MedCheck not be used for?
You want model rankings. It has none.
How are MedCheck scores reported?
Percent of checklist criteria met, per phase. Scores from different benchmarks are not comparable to each other.
Is MedCheck independent?
No conflicts of interest have been verified for MedCheck in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical p
