Radiology visual question answering, with questions written by clinicians about real radiology images.

VQA-RAD is a benchmark in healthcare AI. Radiology visual question answering, with questions written by clinicians about real radiology images. Scores are reported as: Percent accuracy, open and closed questions.

Audited by MedCheckLau et al., Scientific Data 2018

Key facts

What it measuresRadiology visual question answering, with questions written by clinicians about real radiology images.
Who should careRadiology AI teams.
Use it whenClinician-written radiology VQA evaluation.
Do not use it forLarge-scale evaluation. It is small.
Score scalePercent accuracy, open and closed questions.
Grader methodunknown
First released2018
LanguagesEnglish
Use casesimaging
CitationLau et al., Scientific Data 2018
Statusactive

Frequently asked questions

What does VQA-RAD measure?

Radiology visual question answering, with questions written by clinicians about real radiology images.

Who should use VQA-RAD?

Radiology AI teams. Clinician-written radiology VQA evaluation.

What should VQA-RAD not be used for?

Large-scale evaluation. It is small.

How are VQA-RAD scores reported?

Percent accuracy, open and closed questions. Scores from different benchmarks are not comparable to each other.

Is VQA-RAD independent?

No conflicts of interest have been verified for VQA-RAD in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities