The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

MMMU (Health and Medicine) is a benchmark in healthcare AI. The health and medicine subset of a broad expert-level multimodal reasoning benchmark. Scores are reported as: Percent accuracy on the health and medicine subset.

Audited by MedCheckYue et al., 2024

Key facts

What it measuresThe health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Who should careTeams comparing medical performance against general multimodal capability.
Use it whenPlacing medical performance in the context of general reasoning.
Do not use it forDeep clinical evaluation. It is a subset of a general benchmark.
Score scalePercent accuracy on the health and medicine subset.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesimaging, multimodal, exam knowledge
CitationYue et al., 2024
Statusactive

Paper

Frequently asked questions

What does MMMU (Health and Medicine) measure?

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

Who should use MMMU (Health and Medicine)?

Teams comparing medical performance against general multimodal capability. Placing medical performance in the context of general reasoning.

What should MMMU (Health and Medicine) not be used for?

Deep clinical evaluation. It is a subset of a general benchmark.

How are MMMU (Health and Medicine) scores reported?

Percent accuracy on the health and medicine subset. Scores from different benchmarks are not comparable to each other.

Is MMMU (Health and Medicine) independent?

No conflicts of interest have been verified for MMMU (Health and Medicine) in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

PathMMU

Expert-level pathology understanding and reasoning from slide images.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities