Chinese-language evaluation of physical and mental health knowledge in large models.

CHBench is a benchmark in healthcare AI. Chinese-language evaluation of physical and mental health knowledge in large models. Scores are reported as: Percent accuracy.

Audited by MedCheckGuo et al., 2024

Key facts

What it measuresChinese-language evaluation of physical and mental health knowledge in large models.
Who should careTeams deploying in Chinese-speaking markets.
Use it whenEvaluating Chinese health language capability.
Do not use it forDrawing conclusions about English clinical performance.
Score scalePercent accuracy.
Grader methodunknown
First released2024
LanguagesChinese
Use casesexam knowledge, patient communication
CitationGuo et al., 2024
Statusactive

Paper

Frequently asked questions

What does CHBench measure?

Chinese-language evaluation of physical and mental health knowledge in large models.

Who should use CHBench?

Teams deploying in Chinese-speaking markets. Evaluating Chinese health language capability.

What should CHBench not be used for?

Drawing conclusions about English clinical performance.

How are CHBench scores reported?

Percent accuracy. Scores from different benchmarks are not comparable to each other.

Is CHBench independent?

No conflicts of interest have been verified for CHBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AfriMed-QA

Multi-specialty medical questions written by clinicians and students across African countries, built to test w

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

CMExam

Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.

COGNET-MD

Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic

HeadQA

Spanish healthcare specialization exam questions requiring multi-step reasoning.

MedMCQA

Indian medical entrance exam questions across many subjects. Very large and very widely used.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities