Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

CliMedBench is a benchmark in healthcare AI. Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions. Scores are reported as: Scenario-level accuracy and generation metrics.

Audited by MedCheckOuyang et al., EMNLP 2024

Key facts

What it measuresLarge-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Who should careTeams deploying clinical AI in China.
Use it whenChinese clinical scenario evaluation.
Do not use it forEnglish-language deployment decisions.
Score scaleScenario-level accuracy and generation metrics.
Grader methodunknown
First released2024
LanguagesChinese
Use casesclinical decision support
CitationOuyang et al., EMNLP 2024
Statusactive

Frequently asked questions

What does CliMedBench measure?

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

Who should use CliMedBench?

Teams deploying clinical AI in China. Chinese clinical scenario evaluation.

What should CliMedBench not be used for?

English-language deployment decisions.

How are CliMedBench scores reported?

Scenario-level accuracy and generation metrics. Scores from different benchmarks are not comparable to each other.

Is CliMedBench independent?

No conflicts of interest have been verified for CliMedBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities