Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion.

MedHELM is a benchmark in healthcare AI. Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion. Scores are reported as: Mean win rate, 0 to 1. Not comparable to any percent accuracy or Elo figure.

Bedi et al., Nature Medicine 2026Leaderboard 84d old41.9% of datasets public4 verified conflicts

Key facts

What it measuresTests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion.
Who should careHealth system data science teams choosing a model for clinical workflows.
Use it whenYou need task-level rather than exam-level evidence, or you want to run an evaluation on your own patient data inside your network.
Do not use it forYou need current model rankings. The published board mixes model generations and is months behind.
Score scaleMean win rate, 0 to 1. Not comparable to any percent accuracy or Elo figure.
Grader methodICC 0.47 vs expert clinicians
Tasks121
Datasets31 total: 13 public, 6 gated, 12 private
Current leaderGemini 3.1 Pro (Preview) (Google), 0.652
Leaderboard statev5.0.0, updated 2026-05-14, 11 models
First released2025
LanguagesEnglish
Use casesclinical decision support, documentation, patient communication, research, administration
LicenceApache-2.0
CitationBedi et al., Nature Medicine 2026
Statusactive

Official site · Leaderboard · Code · DOI

Who is behind MedHELM

creator
Stanford Center for Research on Foundation Models
Source, verified
creator and commercial
Microsoft Healthcare and Life Sciences
Model vendor co-originated the benchmark.
Source, verified
creator
Stanford Health Care Technology and Digital Solutions
Source, verified
steward and commercial
Pacific AI
Sells: Healthcare AI governance platform with automated evaluation pipelines that include MedHELM accuracy benchmarks
Steward of the benchmark also sells the compliance product that runs it, and is a CHAI-certified Assurance Resource Provider.
Source, verified
code host and commercial
Pacific AI
Community-owned Apache 2.0 code hosted in a commercial vendor GitHub organization rather than a foundation or university namespace.
Source, verified

Verified conflicts of interest

Steward sells a governance product that runs this benchmark
Co-created by a model vendor
Steward certified by the coalition that governs its users
Over half the datasets are private or gated

Independence assessment

Self-described as independent and community-led since 2026, but technical stewardship sits with Pacific AI, a commercial healthcare AI governance vendor that sells model testing pipelines including MedHELM benchmarks, and the code repo lives in the Pacific AI GitHub org. Microsoft Healthcare and Life Sciences was a co-originator. Google models hold ranks 1 and 2 on the current board.

Scope

Clinician-validated taxonomy of 5 categories, 22 subcategories, 121 clinical tasks. Built on Stanford CRFM HELM.

Frequently asked questions

What does MedHELM measure?

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge instructions, supporting case discussion.

Who should use MedHELM?

Health system data science teams choosing a model for clinical workflows. You need task-level rather than exam-level evidence, or you want to run an evaluation on your own patient data inside your network.

What should MedHELM not be used for?

You need current model rankings. The published board mixes model generations and is months behind.

How are MedHELM scores reported?

Mean win rate, 0 to 1. Not comparable to any percent accuracy or Elo figure. Scores from different benchmarks are not comparable to each other.

Is MedHELM independent?

Not fully. Steward sells a governance product that runs this benchmark; Co-created by a model vendor; Steward certified by the coalition that governs its users; Over half the datasets are private or gated.

Is the MedHELM leaderboard current?

The published leaderboard was last updated on 2026-05-14, which is 84 days before 2026-08-06.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities