Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field.

cMedQA2 is a benchmark in healthcare AI. Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field. Scores are reported as: Answer selection accuracy.

Audited by MedCheckZhang et al., IEEE Access 2018

Key facts

What it measuresChinese medical question and answer selection built from online health forums. One of the oldest entries in the field.
Who should careResearchers tracing the history of medical NLP evaluation.
Use it whenHistorical comparison and retrieval task evaluation.
Do not use it forAny current model selection decision. It predates large language models.
Score scaleAnswer selection accuracy.
Grader methodunknown
First released2018
LanguagesChinese
Use casesresearch
CitationZhang et al., IEEE Access 2018
Statusactive

Frequently asked questions

What does cMedQA2 measure?

Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field.

Who should use cMedQA2?

Researchers tracing the history of medical NLP evaluation. Historical comparison and retrieval task evaluation.

What should cMedQA2 not be used for?

Any current model selection decision. It predates large language models.

How are cMedQA2 scores reported?

Answer selection accuracy. Scores from different benchmarks are not comparable to each other.

Is cMedQA2 independent?

No conflicts of interest have been verified for cMedQA2 in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CLIMB

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

DataDEL

Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished

HELM

The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedOdyssey

Long-context medical evaluation running up to 200,000 tokens.

MedS-Bench

Evaluates models across eleven clinical task types beyond question answering, including summarization, extract

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities