Chinese online medical question answering drawn from consumer health platforms.

webMedQA is a benchmark in healthcare AI. Chinese online medical question answering drawn from consumer health platforms. Scores are reported as: Answer selection and ranking metrics.

Audited by MedCheckHe et al., BMC Medical Informatics 2019

Key facts

What it measuresChinese online medical question answering drawn from consumer health platforms.
Who should careTeams building consumer health tools in Chinese.
Use it whenConsumer-facing Chinese health language evaluation.
Do not use it forClinical decision support.
Score scaleAnswer selection and ranking metrics.
Grader methodunknown
First released2019
LanguagesChinese
Use casespatient communication
CitationHe et al., BMC Medical Informatics 2019
Statusactive

Frequently asked questions

What does webMedQA measure?

Chinese online medical question answering drawn from consumer health platforms.

Who should use webMedQA?

Teams building consumer health tools in Chinese. Consumer-facing Chinese health language evaluation.

What should webMedQA not be used for?

Clinical decision support.

How are webMedQA scores reported?

Answer selection and ranking metrics. Scores from different benchmarks are not comparable to each other.

Is webMedQA independent?

No conflicts of interest have been verified for webMedQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CHBench

Chinese-language evaluation of physical and mental health knowledge in large models.

HealthBench

Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing

LMArena Medicine and Healthcare

Open public voting on model answers, filtered to a medicine and healthcare occupational category.

MedArena

Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access

MedExQA

Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities