Expert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track alongside the text track.

MedXpertQA is a benchmark in healthcare AI. Expert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track alongside the text track. Scores are reported as: Percent accuracy, text and multimodal tracks reported separately.

Zuo et al., 2025

Key facts

What it measuresExpert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track alongside the text track.
Who should careTeams whose models already score near the ceiling on MedQA and need headroom.
Use it whenMedQA and MedMCQA have saturated and you need a benchmark that still discriminates.
Do not use it forComparability with the older literature. It is intentionally not the same difficulty.
Score scalePercent accuracy, text and multimodal tracks reported separately.
Grader methodmultiple choice
First released2025
LanguagesEnglish
Use casesexam knowledge, clinical decision support, multimodal
CitationZuo et al., 2025
Statusactive

Paper

Frequently asked questions

What does MedXpertQA measure?

Expert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimodal track alongside the text track.

Who should use MedXpertQA?

Teams whose models already score near the ceiling on MedQA and need headroom. MedQA and MedMCQA have saturated and you need a benchmark that still discriminates.

What should MedXpertQA not be used for?

Comparability with the older literature. It is intentionally not the same difficulty.

How are MedXpertQA scores reported?

Percent accuracy, text and multimodal tracks reported separately. Scores from different benchmarks are not comparable to each other.

Is MedXpertQA independent?

No conflicts of interest have been verified for MedXpertQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AfriMed-QA

Multi-specialty medical questions written by clinicians and students across African countries, built to test w

CHBench

Chinese-language evaluation of physical and mental health knowledge in large models.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

CMExam

Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.

COGNET-MD

Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic

HeadQA

Spanish healthcare specialization exam questions requiring multi-step reasoning.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities