Expert-level pathology understanding and reasoning from slide images.
Audited by MedCheckSun et al., ECCV 2024
Key facts
| What it measures | Expert-level pathology understanding and reasoning from slide images. |
| Who should care | Digital pathology teams. |
| Use it when | Pathology-specific model selection. |
| Do not use it for | General medical evaluation. |
| Score scale | Percent accuracy on expert-level pathology items. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | imaging, multimodal |
| Citation | Sun et al., ECCV 2024 |
| Status | active |
Frequently asked questions
What does PathMMU measure?
Expert-level pathology understanding and reasoning from slide images.
Who should use PathMMU?
Digital pathology teams. Pathology-specific model selection.
What should PathMMU not be used for?
General medical evaluation.
How are PathMMU scores reported?
Percent accuracy on expert-level pathology items. Scores from different benchmarks are not comparable to each other.
Is PathMMU independent?
No conflicts of interest have been verified for PathMMU in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
