Questions answered from real discharge summaries, written by clinicians for real-world practice.

EHRNoteQA is a benchmark in healthcare AI. Questions answered from real discharge summaries, written by clinicians for real-world practice. Scores are reported as: Accuracy against clinician-written reference answers.

Audited by MedCheckKweon et al., NeurIPS 2024

Key facts

What it measuresQuestions answered from real discharge summaries, written by clinicians for real-world practice.
Who should careHealth systems evaluating record summarization and chart question answering.
Use it whenTesting whether a model can actually read a discharge summary.
Do not use it forOpen-ended conversation quality.
Score scaleAccuracy against clinician-written reference answers.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesdocumentation, clinical decision support
CitationKweon et al., NeurIPS 2024
Statusactive

Frequently asked questions

What does EHRNoteQA measure?

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Who should use EHRNoteQA?

Health systems evaluating record summarization and chart question answering. Testing whether a model can actually read a discharge summary.

What should EHRNoteQA not be used for?

Open-ended conversation quality.

How are EHRNoteQA scores reported?

Accuracy against clinician-written reference answers. Scores from different benchmarks are not comparable to each other.

Is EHRNoteQA independent?

No conflicts of interest have been verified for EHRNoteQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedConceptsQA

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse

MEDEC

Whether a model can find and fix clinical errors already present in a note.

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities