Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answer.

MedExQA is a benchmark in healthcare AI. Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answer. Scores are reported as: Accuracy plus explanation quality scoring.

Audited by MedCheckKim et al., BioNLP 2024

Key facts

What it measuresMedical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answer.
Who should careTeams that need explainable output, not just correct output.
Use it whenEvaluating explanation quality alongside correctness.
Do not use it forSpeed comparisons. Explanation scoring is slower.
Score scaleAccuracy plus explanation quality scoring.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesclinical decision support, patient communication
CitationKim et al., BioNLP 2024
Statusactive

Frequently asked questions

What does MedExQA measure?

Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as the answer.

Who should use MedExQA?

Teams that need explainable output, not just correct output. Evaluating explanation quality alongside correctness.

What should MedExQA not be used for?

Speed comparisons. Explanation scoring is slower.

How are MedExQA scores reported?

Accuracy plus explanation quality scoring. Scores from different benchmarks are not comparable to each other.

Is MedExQA independent?

No conflicts of interest have been verified for MedExQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities