Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
Audited by MedCheckLiu et al., ACL 2025
Key facts
| What it measures | Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels. |
| Who should care | Teams evaluating vision-language models on medical images. |
| Use it when | You need breadth across imaging types rather than one specialty. |
| Do not use it for | Text-only model evaluation. |
| Score scale | Percent accuracy across modality and difficulty strata. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | imaging, multimodal |
| Citation | Liu et al., ACL 2025 |
| Status | active |
Frequently asked questions
What does Asclepius measure?
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
Who should use Asclepius?
Teams evaluating vision-language models on medical images. You need breadth across imaging types rather than one specialty.
What should Asclepius not be used for?
Text-only model evaluation.
How are Asclepius scores reported?
Percent accuracy across modality and difficulty strata. Scores from different benchmarks are not comparable to each other.
Is Asclepius independent?
No conflicts of interest have been verified for Asclepius in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
Expert-level pathology understanding and reasoning from slide images.
