Very large medical visual question answering set spanning many modalities and anatomical regions.

OmniMedVQA is a benchmark in healthcare AI. Very large medical visual question answering set spanning many modalities and anatomical regions. Scores are reported as: Percent accuracy across modalities.

Audited by MedCheckHu et al., 2024

Key facts

What it measuresVery large medical visual question answering set spanning many modalities and anatomical regions.
Who should careVision-language model teams needing scale.
Use it whenLarge-scale multimodal medical evaluation.
Do not use it forText-only tasks.
Score scalePercent accuracy across modalities.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesimaging, multimodal
CitationHu et al., 2024
Statusactive

Paper

Frequently asked questions

What does OmniMedVQA measure?

Very large medical visual question answering set spanning many modalities and anatomical regions.

Who should use OmniMedVQA?

Vision-language model teams needing scale. Large-scale multimodal medical evaluation.

What should OmniMedVQA not be used for?

Text-only tasks.

How are OmniMedVQA scores reported?

Percent accuracy across modalities. Scores from different benchmarks are not comparable to each other.

Is OmniMedVQA independent?

No conflicts of interest have been verified for OmniMedVQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

PathMMU

Expert-level pathology understanding and reasoning from slide images.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities