Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Audited by MedCheckLiu et al., 2024
Key facts
| What it measures | Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite. |
| Who should care | Teams wanting one suite across several clinical task families. |
| Use it when | Getting a wide first read across clinical tasks. |
| Do not use it for | Deep evaluation of any single task. Breadth costs depth. |
| Score scale | Mixed per-task metrics. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | clinical decision support, documentation |
| Citation | Liu et al., 2024 |
| Status | active |
Frequently asked questions
What does ClinicBench measure?
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Who should use ClinicBench?
Teams wanting one suite across several clinical task families. Getting a wide first read across clinical tasks.
What should ClinicBench not be used for?
Deep evaluation of any single task. Breadth costs depth.
How are ClinicBench scores reported?
Mixed per-task metrics. Scores from different benchmarks are not comparable to each other.
Is ClinicBench independent?
No conflicts of interest have been verified for ClinicBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
Clinical reasoning in the emergency room, built from real de-identified ER records.
