Whether a model can find and fix clinical errors already present in a note.

MEDEC is a benchmark in healthcare AI. Whether a model can find and fix clinical errors already present in a note. Scores are reported as: Detection and correction accuracy.

Audited by MedCheckBen Abacha et al., 2025

Key facts

What it measuresWhether a model can find and fix clinical errors already present in a note.
Who should careDocumentation quality and scribe review teams.
Use it whenTesting error detection, which is different from error avoidance.
Do not use it forGenerating notes from scratch.
Score scaleDetection and correction accuracy.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesdocumentation, safety
CitationBen Abacha et al., 2025
Statusactive

Paper

Frequently asked questions

What does MEDEC measure?

Whether a model can find and fix clinical errors already present in a note.

Who should use MEDEC?

Documentation quality and scribe review teams. Testing error detection, which is different from error avoidance.

What should MEDEC not be used for?

Generating notes from scratch.

How are MEDEC scores reported?

Detection and correction accuracy. Scores from different benchmarks are not comparable to each other.

Is MEDEC independent?

No conflicts of interest have been verified for MEDEC in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedConceptsQA

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities