Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Audited by MedCheckWang et al., NAACL 2024
Key facts
| What it measures | Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work. |
| Who should care | Teams working in Chinese healthcare. |
| Use it when | Broad Chinese medical capability assessment. |
| Do not use it for | Cross-language comparison. |
| Score scale | Percent accuracy plus expert scoring on open questions. |
| Grader method | unknown |
| First released | 2024 |
| Languages | Chinese |
| Use cases | exam knowledge, clinical decision support |
| Citation | Wang et al., NAACL 2024 |
| Status | active |
Frequently asked questions
What does CMB measure?
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Who should use CMB?
Teams working in Chinese healthcare. Broad Chinese medical capability assessment.
What should CMB not be used for?
Cross-language comparison.
How are CMB scores reported?
Percent accuracy plus expert scoring on open questions. Scores from different benchmarks are not comparable to each other.
Is CMB independent?
No conflicts of interest have been verified for CMB in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Multi-specialty medical questions written by clinicians and students across African countries, built to test w
Chinese-language evaluation of physical and mental health knowledge in large models.
Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.
Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic
Spanish healthcare specialization exam questions requiring multi-step reasoning.
Indian medical entrance exam questions across many subjects. Very large and very widely used.
