Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

Asclepius is a benchmark in healthcare AI. Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels. Scores are reported as: Percent accuracy across modality and difficulty strata.

Audited by MedCheckLiu et al., ACL 2025

Key facts

What it measuresSpectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
Who should careTeams evaluating vision-language models on medical images.
Use it whenYou need breadth across imaging types rather than one specialty.
Do not use it forText-only model evaluation.
Score scalePercent accuracy across modality and difficulty strata.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesimaging, multimodal
CitationLiu et al., ACL 2025
Statusactive

Frequently asked questions

What does Asclepius measure?

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

Who should use Asclepius?

Teams evaluating vision-language models on medical images. You need breadth across imaging types rather than one specialty.

What should Asclepius not be used for?

Text-only model evaluation.

How are Asclepius scores reported?

Percent accuracy across modality and difficulty strata. Scores from different benchmarks are not comparable to each other.

Is Asclepius independent?

No conflicts of interest have been verified for Asclepius in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

PathMMU

Expert-level pathology understanding and reasoning from slide images.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities