Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

CLIMB is a benchmark in healthcare AI. Data foundations for large-scale multimodal clinical models, spanning many data types rather than one. Scores are reported as: Task-specific metrics across the included datasets.

Audited by MedCheckDai et al., ICML 2025

Key facts

What it measuresData foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Who should careTeams pretraining or fine-tuning clinical foundation models.
Use it whenAssessing coverage across clinical data modalities.
Do not use it forRanking off-the-shelf chat models.
Score scaleTask-specific metrics across the included datasets.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesmultimodal, research
CitationDai et al., ICML 2025
Statusactive

Frequently asked questions

What does CLIMB measure?

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

Who should use CLIMB?

Teams pretraining or fine-tuning clinical foundation models. Assessing coverage across clinical data modalities.

What should CLIMB not be used for?

Ranking off-the-shelf chat models.

How are CLIMB scores reported?

Task-specific metrics across the included datasets. Scores from different benchmarks are not comparable to each other.

Is CLIMB independent?

No conflicts of interest have been verified for CLIMB in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MedXpertQA

Expert-level medical reasoning questions deliberately built to be harder than licensing exams, with a multimod

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities