The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline.
Liang et al., TMLR 2023
Key facts
| What it measures | The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline. |
| Who should care | Anyone who wants to understand where the medical benchmark infrastructure came from. |
| Use it when | General model comparison, or as the engine for running a medical suite yourself. |
| Do not use it for | Medical decisions on its own. The general framework is not clinically grounded. |
| Score scale | Per-scenario metrics with a mean win rate rollup. |
| Grader method | mixed |
| First released | 2022 |
| Languages | English |
| Use cases | research |
| Licence | Apache-2.0 |
| Citation | Liang et al., TMLR 2023 |
| Status | active |
Official site · Code · Paper
Who is behind HELM
Independence assessment
Stanford CRFM maintained. Evaluates predefined models chosen by the maintainers rather than accepting submissions.
Scope
Parent framework that MedHELM extends. Stanford CRFM. Not open to public submissions.
Frequently asked questions
What does HELM measure?
The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardized, reproducible pipeline.
Who should use HELM?
Anyone who wants to understand where the medical benchmark infrastructure came from. General model comparison, or as the engine for running a medical suite yourself.
What should HELM not be used for?
Medical decisions on its own. The general framework is not clinically grounded.
How are HELM scores reported?
Per-scenario metrics with a mean win rate rollup. Scores from different benchmarks are not comparable to each other.
Is HELM independent?
No conflicts of interest have been verified for HELM in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Chinese medical question and answer selection built from online health forums. One of the oldest entries in th
Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Long-context medical evaluation running up to 200,000 tokens.
Evaluates models across eleven clinical task types beyond question answering, including summarization, extract
