Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

ClinicBench is a benchmark in healthcare AI. Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite. Scores are reported as: Mixed per-task metrics.

Audited by MedCheckLiu et al., 2024

Key facts

What it measuresBroad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Who should careTeams wanting one suite across several clinical task families.
Use it whenGetting a wide first read across clinical tasks.
Do not use it forDeep evaluation of any single task. Breadth costs depth.
Score scaleMixed per-task metrics.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesclinical decision support, documentation
CitationLiu et al., 2024
Statusactive

Paper

Frequently asked questions

What does ClinicBench measure?

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

Who should use ClinicBench?

Teams wanting one suite across several clinical task families. Getting a wide first read across clinical tasks.

What should ClinicBench not be used for?

Deep evaluation of any single task. Breadth costs depth.

How are ClinicBench scores reported?

Mixed per-task metrics. Scores from different benchmarks are not comparable to each other.

Is ClinicBench independent?

No conflicts of interest have been verified for ClinicBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities