Medical visual question answering from the ImageCLEF challenge series.
Audited by MedCheckBen Abacha et al., CLEF 2021
Key facts
| What it measures | Medical visual question answering from the ImageCLEF challenge series. |
| Who should care | Researchers citing the medical VQA challenge literature. |
| Use it when | Historical comparison. |
| Do not use it for | Current clinical deployment decisions. |
| Score scale | Percent accuracy and BLEU on generated answers. |
| Grader method | unknown |
| First released | 2021 |
| Languages | English |
| Use cases | imaging |
| Citation | Ben Abacha et al., CLEF 2021 |
| Status | active |
Frequently asked questions
What does VQA-Med measure?
Medical visual question answering from the ImageCLEF challenge series.
Who should use VQA-Med?
Researchers citing the medical VQA challenge literature. Historical comparison.
What should VQA-Med not be used for?
Current clinical deployment decisions.
How are VQA-Med scores reported?
Percent accuracy and BLEU on generated answers. Scores from different benchmarks are not comparable to each other.
Is VQA-Med independent?
No conflicts of interest have been verified for VQA-Med in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
