Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.

GMAI-MMBench is a benchmark in healthcare AI. Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types. Scores are reported as: Percent accuracy across modality and department strata.

Audited by MedCheckChen et al., NeurIPS 2024

Key facts

What it measuresLarge multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.
Who should careTeams evaluating vision-language models broadly across medicine.
Use it whenYou need one wide multimodal read.
Do not use it forText-only evaluation, or depth in one specialty.
Score scalePercent accuracy across modality and department strata.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesimaging, multimodal
CitationChen et al., NeurIPS 2024
Statusactive

Frequently asked questions

What does GMAI-MMBench measure?

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.

Who should use GMAI-MMBench?

Teams evaluating vision-language models broadly across medicine. You need one wide multimodal read.

What should GMAI-MMBench not be used for?

Text-only evaluation, or depth in one specialty.

How are GMAI-MMBench scores reported?

Percent accuracy across modality and department strata. Scores from different benchmarks are not comparable to each other.

Is GMAI-MMBench independent?

No conflicts of interest have been verified for GMAI-MMBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

PathMMU

Expert-level pathology understanding and reasoning from slide images.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities