Who grades the AI doctors, who owns the scoreboard, and how old the numbers are.

The Healthcare AI Benchmark Registry catalogues 63 evaluation instruments used to score medical AI, covering 56 benchmarks, 2 arenas, and the audit frameworks and certification bodies around them. For each one it records what it actually measures, who created and funds it, who stewards the code, whether that steward sells a product depending on it, what scale the scores use, and how many days old the published leaderboard is. Scores from different benchmarks are not comparable to each other, which is the single most common error in healthcare AI procurement.

63 entities63 fully profiled6 with verified conflicts56 scoring criteriaUpdated 2026-08-05

How to use this registry

Start from what you are deploying, not from the rankings. If you are choosing a model for clinical documentation, the exam benchmarks tell you almost nothing. If you are launching a patient-facing tool, a safety benchmark matters more than any accuracy score. Each entry states plainly what it is for and what it should not be used for.

Then check two things most catalogues omit. First, who owns the scoreboard, because several major benchmarks are created by model vendors or stewarded by companies that sell products depending on them. Second, how old the published numbers are, because a leaderboard months behind the current model generation will rank a model you cannot buy above one you can.

All 63 entities

NameTypeWhat it measures
AfriMed-QAbenchmarkMulti-specialty medical questions written by clinicians and students across African countries, built to test whether mod
AgentClinicbenchmarkSimulated clinical encounters where the model must gather information over several turns rather than answer one question
AsclepiusbenchmarkSpectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
Awesome Agents Medical LLM LeaderboardaggregatorAggregates published medical benchmark scores into one ranking table.
BetterBenchaudit frameworkAssesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical predecessor
BRIDGEbenchmarkReal clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.
CHAIcertification bodyAn industry coalition that publishes model cards and certifies the vendors who audit healthcare AI. It certifies assuran
CHBenchbenchmarkChinese-language evaluation of physical and mental health knowledge in large models.
CheXpertbenchmarkA large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image
CLIMBbenchmarkData foundations for large-scale multimodal clinical models, spanning many data types rather than one.
CliMedBenchbenchmarkLarge-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
ClinicBenchbenchmarkBroad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
CMBbenchmarkComprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
cMedQA2benchmarkChinese medical question and answer selection built from online health forums. One of the oldest entries in the field.
CMExambenchmarkChinese medical licensing exam questions with expert annotations for reasoning and difficulty.
COGNET-MDbenchmarkEvaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic questions
DataDELbenchmarkMeasures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answ
EHRNoteQAbenchmarkQuestions answered from real discharge summaries, written by clinicians for real-world practice.
EndoBenchbenchmarkMultimodal evaluation for endoscopy image and video analysis.
ER-REASONbenchmarkClinical reasoning in the emergency room, built from real de-identified ER records.
GMAI-MMBenchbenchmarkLarge multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.
HeadQAbenchmarkSpanish healthcare specialization exam questions requiring multi-step reasoning.
HealthBenchbenchmarkPhysician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.
HELMbenchmarkThe parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reprod
LLMEval-MedbenchmarkReal-world clinical benchmark with physician validation of both questions and reference answers.
LLM-Stats Healthcare LeaderboardaggregatorRanks models on health using its own composite index.
LMArena Medicine and HealthcarearenaOpen public voting on model answers, filtered to a medicine and healthcare occupational category.
MedAgentBenchbenchmarkA realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answeri
MedAgentsBenchbenchmarkCompares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fails.
MedArenaarenaPracticing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restric
MedCalc-BenchbenchmarkWhether a model computes clinical scores and formulas correctly from a patient note.
MedChainbenchmarkInteractive sequential benchmarking that follows a case through stages the way a real clinic does.
MedCheckaudit frameworkAudits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.
MedConceptsQAbenchmarkWhether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.
MEDECbenchmarkWhether a model can find and fix clinical errors already present in a note.
MedExQAbenchmarkMedical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answe
MedHELMbenchmarkTests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions,
MEDICbenchmarkA framework covering several clinical application dimensions at once, including safety and reasoning, rather than a sing
MediQbenchmarkWhether a model asks the right follow-up question when it does not have enough information, instead of guessing.
MedJourneybenchmarkFollows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated momen
MedMCQAbenchmarkIndian medical entrance exam questions across many subjects. Very large and very widely used.
MedOdysseybenchmarkLong-context medical evaluation running up to 200,000 tokens.
MedQAbenchmarkUS medical licensing exam style questions. The original medical LLM benchmark and still the most cited.
MedR-BenchbenchmarkQuantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer.
MedRiskbenchmarkGeneralist clinical risk prediction across many validated risk tools, run as an autonomous copilot task.
MedSafetyBenchbenchmarkWhether a model refuses or complies with medically harmful requests.
MedS-BenchbenchmarkEvaluates models across eleven clinical task types beyond question answering, including summarization, extraction and ex
MedXpertQAbenchmarkExpert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track a
MLEC-QAbenchmarkChinese medical licensing exam multiple choice across five subfields.
MMMU (Health and Medicine)benchmarkThe health and medicine subset of a broad expert-level multimodal reasoning benchmark.
MVME (AI Hospital)benchmarkA multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
OmniMedVQAbenchmarkVery large medical visual question answering set spanning many modalities and anatomical regions.
PathMMUbenchmarkExpert-level pathology understanding and reasoning from slide images.
PathVQAbenchmarkPathology image question answering, one of the earliest medical VQA sets.
PMC-VQAbenchmarkMedical visual question answering built from PubMed Central figures at large scale.
PubMedQAbenchmarkYes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.
ReasonMedbenchmarkA very large multi-agent generated reasoning dataset for medical training and evaluation.
SLAKEbenchmarkBilingual medical visual question answering with a semantic knowledge layer attached to the images.
TrialPanoramabenchmarkDatabase and benchmark for systematic review work and clinical trial design.
VQA-MedbenchmarkMedical visual question answering from the ImageCLEF challenge series.
VQA-RADbenchmarkRadiology visual question answering, with questions written by clinicians about real radiology images.
webMedQAbenchmarkChinese online medical question answering drawn from consumer health platforms.
XMedBenchbenchmarkMultilingual medical benchmark covering six widely spoken languages, built to widen access beyond English.

Frequently asked questions

What is a healthcare AI benchmark?

A healthcare AI benchmark is a standardized set of medical tasks used to score how well a language model performs. Some test exam knowledge, some test real clinical workflows such as summarizing a record or drafting discharge instructions, and some collect clinician preferences instead of measuring accuracy.

Which healthcare AI benchmark is best?

There is no single best one, because they measure different things on different scales. Match the benchmark to the deployment: exam benchmarks for knowledge comparison, task benchmarks for clinical workflows, safety benchmarks before a patient-facing launch, and agentic benchmarks for tools that take actions.

Can benchmark scores be compared across leaderboards?

No. A mean win rate, an Elo rating and a percent accuracy are different units. Two leaderboards can rank the same models in opposite orders without either being wrong. Every entity page in this registry states its own scale for that reason.

Who owns the healthcare AI benchmarks?

Ownership is mixed and often undisclosed. Some are created by model vendors who also report their own results. Some are stewarded by commercial governance companies that sell products depending on the benchmark. This registry records the creator, funder, steward and code host for each entity, with a linked source.

How current are healthcare AI leaderboards?

Frequently months out of date. Leaderboards publish current state only and discard history, so staleness is invisible unless someone tracks it. This registry records the last published update date and computes the age in days.

What is MedCheck?

MedCheck is a peer-reviewed lifecycle audit framework that scored medical benchmarks against 46 criteria across five phases, from design through governance. Its inventory is the seed for this registry. It audits benchmark quality but does not record ownership or funding, which is the layer added here.

Methodology in brief

The seed inventory comes from the MedCheck audit, a peer-reviewed lifecycle assessment of medical benchmarks published at ACL 2026. We adopted its 46 criteria verbatim as the base instrument and added a sixth phase covering ownership, incentives and independence, which MedCheck does not measure.

Every governance claim on this site carries a linked primary source and a verification date. Claims are labelled verified, supported, disputed or unknown. We publish no single independence score, because a composite we have not validated would be the same failure this registry documents in others. Full detail is on the methodology page.

Compiled and maintained by Healthcare Discovery. Last updated 2026-08-05. Corrections welcome.