Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Audited by MedCheckOuyang et al., EMNLP 2024
Key facts
| What it measures | Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions. |
| Who should care | Teams deploying clinical AI in China. |
| Use it when | Chinese clinical scenario evaluation. |
| Do not use it for | English-language deployment decisions. |
| Score scale | Scenario-level accuracy and generation metrics. |
| Grader method | unknown |
| First released | 2024 |
| Languages | Chinese |
| Use cases | clinical decision support |
| Citation | Ouyang et al., EMNLP 2024 |
| Status | active |
Frequently asked questions
What does CliMedBench measure?
Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Who should use CliMedBench?
Teams deploying clinical AI in China. Chinese clinical scenario evaluation.
What should CliMedBench not be used for?
English-language deployment decisions.
How are CliMedBench scores reported?
Scenario-level accuracy and generation metrics. Scores from different benchmarks are not comparable to each other.
Is CliMedBench independent?
No conflicts of interest have been verified for CliMedBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
Clinical reasoning in the emergency room, built from real de-identified ER records.
