Radiology visual question answering, with questions written by clinicians about real radiology images.
Audited by MedCheckLau et al., Scientific Data 2018
Key facts
| What it measures | Radiology visual question answering, with questions written by clinicians about real radiology images. |
| Who should care | Radiology AI teams. |
| Use it when | Clinician-written radiology VQA evaluation. |
| Do not use it for | Large-scale evaluation. It is small. |
| Score scale | Percent accuracy, open and closed questions. |
| Grader method | unknown |
| First released | 2018 |
| Languages | English |
| Use cases | imaging |
| Citation | Lau et al., Scientific Data 2018 |
| Status | active |
Frequently asked questions
What does VQA-RAD measure?
Radiology visual question answering, with questions written by clinicians about real radiology images.
Who should use VQA-RAD?
Radiology AI teams. Clinician-written radiology VQA evaluation.
What should VQA-RAD not be used for?
Large-scale evaluation. It is small.
How are VQA-RAD scores reported?
Percent accuracy, open and closed questions. Scores from different benchmarks are not comparable to each other.
Is VQA-RAD independent?
No conflicts of interest have been verified for VQA-RAD in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
