Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field.
Audited by MedCheckZhang et al., IEEE Access 2018
Key facts
| What it measures | Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field. |
| Who should care | Researchers tracing the history of medical NLP evaluation. |
| Use it when | Historical comparison and retrieval task evaluation. |
| Do not use it for | Any current model selection decision. It predates large language models. |
| Score scale | Answer selection accuracy. |
| Grader method | unknown |
| First released | 2018 |
| Languages | Chinese |
| Use cases | research |
| Citation | Zhang et al., IEEE Access 2018 |
| Status | active |
Frequently asked questions
What does cMedQA2 measure?
Chinese medical question and answer selection built from online health forums. One of the oldest entries in the field.
Who should use cMedQA2?
Researchers tracing the history of medical NLP evaluation. Historical comparison and retrieval task evaluation.
What should cMedQA2 not be used for?
Any current model selection decision. It predates large language models.
How are cMedQA2 scores reported?
Answer selection accuracy. Scores from different benchmarks are not comparable to each other.
Is cMedQA2 independent?
No conflicts of interest have been verified for cMedQA2 in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished
The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Long-context medical evaluation running up to 200,000 tokens.
Evaluates models across eleven clinical task types beyond question answering, including summarization, extract
