Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.
Audited by MedCheckChen et al., NeurIPS 2024
Key facts
| What it measures | Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types. |
| Who should care | Teams evaluating vision-language models broadly across medicine. |
| Use it when | You need one wide multimodal read. |
| Do not use it for | Text-only evaluation, or depth in one specialty. |
| Score scale | Percent accuracy across modality and department strata. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | imaging, multimodal |
| Citation | Chen et al., NeurIPS 2024 |
| Status | active |
Frequently asked questions
What does GMAI-MMBench measure?
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task types.
Who should use GMAI-MMBench?
Teams evaluating vision-language models broadly across medicine. You need one wide multimodal read.
What should GMAI-MMBench not be used for?
Text-only evaluation, or depth in one specialty.
How are GMAI-MMBench scores reported?
Percent accuracy across modality and department strata. Scores from different benchmarks are not comparable to each other.
Is GMAI-MMBench independent?
No conflicts of interest have been verified for GMAI-MMBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
Expert-level pathology understanding and reasoning from slide images.
