Whether a model can find and fix clinical errors already present in a note.
Audited by MedCheckBen Abacha et al., 2025
Key facts
| What it measures | Whether a model can find and fix clinical errors already present in a note. |
| Who should care | Documentation quality and scribe review teams. |
| Use it when | Testing error detection, which is different from error avoidance. |
| Do not use it for | Generating notes from scratch. |
| Score scale | Detection and correction accuracy. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | documentation, safety |
| Citation | Ben Abacha et al., 2025 |
| Status | active |
Frequently asked questions
What does MEDEC measure?
Whether a model can find and fix clinical errors already present in a note.
Who should use MEDEC?
Documentation quality and scribe review teams. Testing error detection, which is different from error avoidance.
What should MEDEC not be used for?
Generating notes from scratch.
How are MEDEC scores reported?
Detection and correction accuracy. Scores from different benchmarks are not comparable to each other.
Is MEDEC independent?
No conflicts of interest have been verified for MEDEC in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
