Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation.

MedS-Bench is a benchmark in healthcare AI. Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation. Scores are reported as: Per-task metrics across the eleven types.

Audited by MedCheckWu et al., npj Digital Medicine 2025

Key facts

What it measuresEvaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation.
Who should careTeams needing coverage of the tasks that actually fill a clinical day.
Use it whenBroad task-type coverage in one suite.
Do not use it forDepth in any single task.
Score scalePer-task metrics across the eleven types.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesdocumentation, clinical decision support, research
CitationWu et al., npj Digital Medicine 2025
Statusactive

Frequently asked questions

What does MedS-Bench measure?

Evaluates models across eleven clinical task types beyond question answering, including summarization, extraction and explanation.

Who should use MedS-Bench?

Teams needing coverage of the tasks that actually fill a clinical day. Broad task-type coverage in one suite.

What should MedS-Bench not be used for?

Depth in any single task.

How are MedS-Bench scores reported?

Per-task metrics across the eleven types. Scores from different benchmarks are not comparable to each other.

Is MedS-Bench independent?

No conflicts of interest have been verified for MedS-Bench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedConceptsQA

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse

MEDEC

Whether a model can find and fix clinical errors already present in a note.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities