Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion.
Bedi et al., Nature Medicine 2026Leaderboard 84d old41.9% of datasets public4 verified conflicts
Key facts
| What it measures | Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion. |
| Who should care | Health system data science teams choosing a model for clinical workflows. |
| Use it when | You need task-level rather than exam-level evidence, or you want to run an evaluation on your own patient data inside your network. |
| Do not use it for | You need current model rankings. The published board mixes model generations and is months behind. |
| Score scale | Mean win rate, 0 to 1. Not comparable to any percent accuracy or Elo figure. |
| Grader method | ICC 0.47 vs expert clinicians |
| Tasks | 121 |
| Datasets | 31 total: 13 public, 6 gated, 12 private |
| Current leader | Gemini 3.1 Pro (Preview) (Google), 0.652 |
| Leaderboard state | v5.0.0, updated 2026-05-14, 11 models |
| First released | 2025 |
| Languages | English |
| Use cases | clinical decision support, documentation, patient communication, research, administration |
| Licence | Apache-2.0 |
| Citation | Bedi et al., Nature Medicine 2026 |
| Status | active |
Official site · Leaderboard · Code · DOI
Who is behind MedHELM
Verified conflicts of interest
Independence assessment
Self-described as independent and community-led since 2026, but technical stewardship sits with Pacific AI, a commercial healthcare AI governance vendor that sells model testing pipelines including MedHELM benchmarks, and the code repo lives in the Pacific AI GitHub org. Microsoft Healthcare and Life Sciences was a co-originator. Google models hold ranks 1 and 2 on the current board.
Scope
Clinician-validated taxonomy of 5 categories, 22 subcategories, 121 clinical tasks. Built on Stanford CRFM HELM.
Frequently asked questions
What does MedHELM measure?
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion.
Who should use MedHELM?
Health system data science teams choosing a model for clinical workflows. You need task-level rather than exam-level evidence, or you want to run an evaluation on your own patient data inside your network.
What should MedHELM not be used for?
You need current model rankings. The published board mixes model generations and is months behind.
How are MedHELM scores reported?
Mean win rate, 0 to 1. Not comparable to any percent accuracy or Elo figure. Scores from different benchmarks are not comparable to each other.
Is MedHELM independent?
Not fully. Steward sells a governance product that runs this benchmark; Co-created by a model vendor; Steward certified by the coalition that governs its users; Over half the datasets are private or gated.
Is the MedHELM leaderboard current?
The published leaderboard was last updated on 2026-05-14, which is 84 days before 2026-08-06.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
