The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline.

HELM is a benchmark in healthcare AI. The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline. Scores are reported as: Per-scenario metrics with a mean win rate rollup.

Liang et al., TMLR 2023

Key facts

What it measuresThe parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline.
Who should careAnyone who wants to understand where the medical benchmark infrastructure came from.
Use it whenGeneral model comparison, or as the engine for running a medical suite yourself.
Do not use it forMedical decisions on its own. The general framework is not clinically grounded.
Score scalePer-scenario metrics with a mean win rate rollup.
Grader methodmixed
First released2022
LanguagesEnglish
Use casesresearch
LicenceApache-2.0
CitationLiang et al., TMLR 2023
Statusactive

Official site · Code · Paper

Who is behind HELM

maintainer
Stanford Center for Research on Foundation Models
Source, verified

Independence assessment

Stanford CRFM maintained. Evaluates predefined models chosen by the maintainers rather than accepting submissions.

Scope

Parent framework that MedHELM extends. Stanford CRFM. Not open to public submissions.

Frequently asked questions

What does HELM measure?

The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline.

Who should use HELM?

Anyone who wants to understand where the medical benchmark infrastructure came from. General model comparison, or as the engine for running a medical suite yourself.

What should HELM not be used for?

Medical decisions on its own. The general framework is not clinically grounded.

How are HELM scores reported?

Per-scenario metrics with a mean win rate rollup. Scores from different benchmarks are not comparable to each other.

Is HELM independent?

No conflicts of interest have been verified for HELM in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CLIMB

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

cMedQA2

Chinese medical question and answer selection built from online health forums. One of the oldest entries in th

DataDEL

Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedOdyssey

Long-context medical evaluation running up to 200,000 tokens.

MedS-Bench

Evaluates models across eleven clinical task types beyond question answering, including summarization, extract

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities