Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation.
Audited by MedCheckWu et al., npj Digital Medicine 2025
Key facts
| What it measures | Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation. |
| Who should care | Teams needing coverage of the tasks that actually fill a clinical day. |
| Use it when | Broad task-type coverage in one suite. |
| Do not use it for | Depth in any single task. |
| Score scale | Per-task metrics across the eleven types. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | documentation, clinical decision support, research |
| Citation | Wu et al., npj Digital Medicine 2025 |
| Status | active |
Frequently asked questions
What does MedS-Bench measure?
Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation.
Who should use MedS-Bench?
Teams needing coverage of the tasks that actually fill a clinical day. Broad task-type coverage in one suite.
What should MedS-Bench not be used for?
Depth in any single task.
How are MedS-Bench scores reported?
Per-task metrics across the eleven types. Scores from different benchmarks are not comparable to each other.
Is MedS-Bench independent?
No conflicts of interest have been verified for MedS-Bench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse
Whether a model can find and fix clinical errors already present in a note.
