Chinese-language evaluation of physical and mental health knowledge in large models.
Audited by MedCheckGuo et al., 2024
Key facts
| What it measures | Chinese-language evaluation of physical and mental health knowledge in large models. |
| Who should care | Teams deploying in Chinese-speaking markets. |
| Use it when | Evaluating Chinese health language capability. |
| Do not use it for | Drawing conclusions about English clinical performance. |
| Score scale | Percent accuracy. |
| Grader method | unknown |
| First released | 2024 |
| Languages | Chinese |
| Use cases | exam knowledge, patient communication |
| Citation | Guo et al., 2024 |
| Status | active |
Frequently asked questions
What does CHBench measure?
Chinese-language evaluation of physical and mental health knowledge in large models.
Who should use CHBench?
Teams deploying in Chinese-speaking markets. Evaluating Chinese health language capability.
What should CHBench not be used for?
Drawing conclusions about English clinical performance.
How are CHBench scores reported?
Percent accuracy. Scores from different benchmarks are not comparable to each other.
Is CHBench independent?
No conflicts of interest have been verified for CHBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Multi-specialty medical questions written by clinicians and students across African countries, built to test w
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.
Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic
Spanish healthcare specialization exam questions requiring multi-step reasoning.
Indian medical entrance exam questions across many subjects. Very large and very widely used.
