Who grades the AI doctors, who owns the scoreboard, and how old the numbers are.
63 entities63 fully profiled6 with verified conflicts56 scoring criteriaUpdated 2026-08-05
How to use this registry
Start from what you are deploying, not from the rankings. If you are choosing a model for clinical documentation, the exam benchmarks tell you almost nothing. If you are launching a patient-facing tool, a safety benchmark matters more than any accuracy score. Each entry states plainly what it is for and what it should not be used for.
Then check two things most catalogues omit. First, who owns the scoreboard, because several major benchmarks are created by model vendors or stewarded by companies that sell products depending on them. Second, how old the published numbers are, because a leaderboard months behind the current model generation will rank a model you cannot buy above one you can.
All 63 entities
| Name | Type | What it measures |
|---|---|---|
| AfriMed-QA | benchmark | Multi-specialty medical questions written by clinicians and students across African countries, built to test whether mod |
| AgentClinic | benchmark | Simulated clinical encounters where the model must gather information over several turns rather than answer one question |
| Asclepius | benchmark | Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels. |
| Awesome Agents Medical LLM Leaderboard | aggregator | Aggregates published medical benchmark scores into one ranking table. |
| BetterBench | audit framework | Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical predecessor |
| BRIDGE | benchmark | Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks. |
| CHAI | certification body | An industry coalition that publishes model cards and certifies the vendors who audit healthcare AI. It certifies assuran |
| CHBench | benchmark | Chinese-language evaluation of physical and mental health knowledge in large models. |
| CheXpert | benchmark | A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image |
| CLIMB | benchmark | Data foundations for large-scale multimodal clinical models, spanning many data types rather than one. |
| CliMedBench | benchmark | Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions. |
| ClinicBench | benchmark | Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite. |
| CMB | benchmark | Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work. |
| cMedQA2 | benchmark | Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field. |
| CMExam | benchmark | Chinese medical licensing exam questions with expert annotations for reasoning and difficulty. |
| COGNET-MD | benchmark | Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic questions |
| DataDEL | benchmark | Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answ |
| EHRNoteQA | benchmark | Questions answered from real discharge summaries, written by clinicians for real-world practice. |
| EndoBench | benchmark | Multimodal evaluation for endoscopy image and video analysis. |
| ER-REASON | benchmark | Clinical reasoning in the emergency room, built from real de-identified ER records. |
| GMAI-MMBench | benchmark | Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types. |
| HeadQA | benchmark | Spanish healthcare specialization exam questions requiring multi-step reasoning. |
| HealthBench | benchmark | Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones. |
| HELM | benchmark | The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reprod |
| LLMEval-Med | benchmark | Real-world clinical benchmark with physician validation of both questions and reference answers. |
| LLM-Stats Healthcare Leaderboard | aggregator | Ranks models on health using its own composite index. |
| LMArena Medicine and Healthcare | arena | Open public voting on model answers, filtered to a medicine and healthcare occupational category. |
| MedAgentBench | benchmark | A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answeri |
| MedAgentsBench | benchmark | Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fails. |
| MedArena | arena | Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restric |
| MedCalc-Bench | benchmark | Whether a model computes clinical scores and formulas correctly from a patient note. |
| MedChain | benchmark | Interactive sequential benchmarking that follows a case through stages the way a real clinic does. |
| MedCheck | audit framework | Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models. |
| MedConceptsQA | benchmark | Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes. |
| MEDEC | benchmark | Whether a model can find and fix clinical errors already present in a note. |
| MedExQA | benchmark | Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answe |
| MedHELM | benchmark | Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, |
| MEDIC | benchmark | A framework covering several clinical application dimensions at once, including safety and reasoning, rather than a sing |
| MediQ | benchmark | Whether a model asks the right follow-up question when it does not have enough information, instead of guessing. |
| MedJourney | benchmark | Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated momen |
| MedMCQA | benchmark | Indian medical entrance exam questions across many subjects. Very large and very widely used. |
| MedOdyssey | benchmark | Long-context medical evaluation running up to 200,000 tokens. |
| MedQA | benchmark | US medical licensing exam style questions. The original medical LLM benchmark and still the most cited. |
| MedR-Bench | benchmark | Quantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer. |
| MedRisk | benchmark | Generalist clinical risk prediction across many validated risk tools, run as an autonomous copilot task. |
| MedSafetyBench | benchmark | Whether a model refuses or complies with medically harmful requests. |
| MedS-Bench | benchmark | Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and ex |
| MedXpertQA | benchmark | Expert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track a |
| MLEC-QA | benchmark | Chinese medical licensing exam multiple choice across five subfields. |
| MMMU (Health and Medicine) | benchmark | The health and medicine subset of a broad expert-level multimodal reasoning benchmark. |
| MVME (AI Hospital) | benchmark | A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction. |
| OmniMedVQA | benchmark | Very large medical visual question answering set spanning many modalities and anatomical regions. |
| PathMMU | benchmark | Expert-level pathology understanding and reasoning from slide images. |
| PathVQA | benchmark | Pathology image question answering, one of the earliest medical VQA sets. |
| PMC-VQA | benchmark | Medical visual question answering built from PubMed Central figures at large scale. |
| PubMedQA | benchmark | Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts. |
| ReasonMed | benchmark | A very large multi-agent generated reasoning dataset for medical training and evaluation. |
| SLAKE | benchmark | Bilingual medical visual question answering with a semantic knowledge layer attached to the images. |
| TrialPanorama | benchmark | Database and benchmark for systematic review work and clinical trial design. |
| VQA-Med | benchmark | Medical visual question answering from the ImageCLEF challenge series. |
| VQA-RAD | benchmark | Radiology visual question answering, with questions written by clinicians about real radiology images. |
| webMedQA | benchmark | Chinese online medical question answering drawn from consumer health platforms. |
| XMedBench | benchmark | Multilingual medical benchmark covering six widely spoken languages, built to widen access beyond English. |
Frequently asked questions
What is a healthcare AI benchmark?
A healthcare AI benchmark is a standardized set of medical tasks used to score how well a language model performs. Some test exam knowledge, some test real clinical workflows such as summarizing a record or drafting discharge instructions, and some collect clinician preferences instead of measuring accuracy.
Which healthcare AI benchmark is best?
There is no single best one, because they measure different things on different scales. Match the benchmark to the deployment: exam benchmarks for knowledge comparison, task benchmarks for clinical workflows, safety benchmarks before a patient-facing launch, and agentic benchmarks for tools that take actions.
Can benchmark scores be compared across leaderboards?
No. A mean win rate, an Elo rating and a percent accuracy are different units. Two leaderboards can rank the same models in opposite orders without either being wrong. Every entity page in this registry states its own scale for that reason.
Who owns the healthcare AI benchmarks?
Ownership is mixed and often undisclosed. Some are created by model vendors who also report their own results. Some are stewarded by commercial governance companies that sell products depending on the benchmark. This registry records the creator, funder, steward and code host for each entity, with a linked source.
How current are healthcare AI leaderboards?
Frequently months out of date. Leaderboards publish current state only and discard history, so staleness is invisible unless someone tracks it. This registry records the last published update date and computes the age in days.
What is MedCheck?
MedCheck is a peer-reviewed lifecycle audit framework that scored medical benchmarks against 46 criteria across five phases, from design through governance. Its inventory is the seed for this registry. It audits benchmark quality but does not record ownership or funding, which is the layer added here.
Methodology in brief
The seed inventory comes from the MedCheck audit, a peer-reviewed lifecycle assessment of medical benchmarks published at ACL 2026. We adopted its 46 criteria verbatim as the base instrument and added a sixth phase covering ownership, incentives and independence, which MedCheck does not measure.
Every governance claim on this site carries a linked primary source and a verification date. Claims are labelled verified, supported, disputed or unknown. We publish no single independence score, because a composite we have not validated would be the same failure this registry documents in others. Full detail is on the methodology page.
