Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.

MedCheck is a audit framework in healthcare AI. Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models. Scores are reported as: Percent of checklist criteria met, per phase.

Ma et al., ACL 2026

Key facts

What it measuresAudits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.
Who should careAnyone deciding which benchmark to trust before they trust its scores.
Use it whenYou want to know whether a benchmark deserves attention at all.
Do not use it forYou want model rankings. It has none.
Score scalePercent of checklist criteria met, per phase.
Grader methodThree independent NLP researchers, discrepancies resolved by consensus
First released2025
LanguagesEnglish
Use casesmeta
CitationMa et al., ACL 2026
Statusactive

Official site · Code · Paper · DOI

Independence assessment

Academic, four universities, no commercial sponsor identified. Audits benchmarks but does not track ownership, funding or steward conflicts, which is the gap this registry fills.

Scope

Five-phase, 46-criteria lifecycle framework applied to 53 medical LLM benchmarks (56 in the ACL version). Base instrument adopted by this registry.

Frequently asked questions

What does MedCheck measure?

Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.

Who should use MedCheck?

Anyone deciding which benchmark to trust before they trust its scores. You want to know whether a benchmark deserves attention at all.

What should MedCheck not be used for?

You want model rankings. It has none.

How are MedCheck scores reported?

Percent of checklist criteria met, per phase. Scores from different benchmarks are not comparable to each other.

Is MedCheck independent?

No conflicts of interest have been verified for MedCheck in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

BetterBench

Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical p

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities