Medical visual question answering built from PubMed Central figures at large scale.
Audited by MedCheckZhang et al., 2024
Key facts
| What it measures | Medical visual question answering built from PubMed Central figures at large scale. |
| Who should care | Vision-language teams needing volume. |
| Use it when | Large-scale medical VQA training and evaluation. |
| Do not use it for | Clinical deployment decisions. Figures are not clinical images. |
| Score scale | Percent accuracy. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | imaging, research |
| Citation | Zhang et al., 2024 |
| Status | active |
Frequently asked questions
What does PMC-VQA measure?
Medical visual question answering built from PubMed Central figures at large scale.
Who should use PMC-VQA?
Vision-language teams needing volume. Large-scale medical VQA training and evaluation.
What should PMC-VQA not be used for?
Clinical deployment decisions. Figures are not clinical images.
How are PMC-VQA scores reported?
Percent accuracy. Scores from different benchmarks are not comparable to each other.
Is PMC-VQA independent?
No conflicts of interest have been verified for PMC-VQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
