A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image classification benchmark.

CheXpert is a benchmark in healthcare AI. A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image classification benchmark. Scores are reported as: AUC and expert radiologist comparison.

Audited by MedCheckIrvin et al., AAAI 2019

Key facts

What it measuresA large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image classification benchmark.
Who should careRadiology AI teams and anyone citing the imaging AI literature.
Use it whenChest imaging model evaluation with an established expert baseline.
Do not use it forEvaluating language models. It is not a text benchmark.
Score scaleAUC and expert radiologist comparison.
Grader methodunknown
First released2019
LanguagesEnglish
Use casesimaging
CitationIrvin et al., AAAI 2019
Statusactive

Frequently asked questions

What does CheXpert measure?

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and is an image classification benchmark.

Who should use CheXpert?

Radiology AI teams and anyone citing the imaging AI literature. Chest imaging model evaluation with an established expert baseline.

What should CheXpert not be used for?

Evaluating language models. It is not a text benchmark.

How are CheXpert scores reported?

AUC and expert radiologist comparison. Scores from different benchmarks are not comparable to each other.

Is CheXpert independent?

No conflicts of interest have been verified for CheXpert in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

PathMMU

Expert-level pathology understanding and reasoning from slide images.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities