Very large medical visual question answering set spanning many modalities and anatomical regions.
Audited by MedCheckHu et al., 2024
Key facts
| What it measures | Very large medical visual question answering set spanning many modalities and anatomical regions. |
| Who should care | Vision-language model teams needing scale. |
| Use it when | Large-scale multimodal medical evaluation. |
| Do not use it for | Text-only tasks. |
| Score scale | Percent accuracy across modalities. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | imaging, multimodal |
| Citation | Hu et al., 2024 |
| Status | active |
Frequently asked questions
What does OmniMedVQA measure?
Very large medical visual question answering set spanning many modalities and anatomical regions.
Who should use OmniMedVQA?
Vision-language model teams needing scale. Large-scale multimodal medical evaluation.
What should OmniMedVQA not be used for?
Text-only tasks.
How are OmniMedVQA scores reported?
Percent accuracy across modalities. Scores from different benchmarks are not comparable to each other.
Is OmniMedVQA independent?
No conflicts of interest have been verified for OmniMedVQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Expert-level pathology understanding and reasoning from slide images.
