Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.

PubMedQA is a benchmark in healthcare AI. Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts. Scores are reported as: Percent accuracy across three answer classes.

Audited by MedCheckJin et al., EMNLP-IJCNLP 2019

Key facts

What it measuresYes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.
Who should careTeams building literature research tools.
Use it whenEvaluating literature comprehension.
Do not use it forEvaluating clinical reasoning.
Score scalePercent accuracy across three answer classes.
Grader methodunknown
First released2019
LanguagesEnglish
Use casesresearch
CitationJin et al., EMNLP-IJCNLP 2019
Statusactive

Frequently asked questions

What does PubMedQA measure?

Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.

Who should use PubMedQA?

Teams building literature research tools. Evaluating literature comprehension.

What should PubMedQA not be used for?

Evaluating clinical reasoning.

How are PubMedQA scores reported?

Percent accuracy across three answer classes. Scores from different benchmarks are not comparable to each other.

Is PubMedQA independent?

No conflicts of interest have been verified for PubMedQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CLIMB

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

cMedQA2

Chinese medical question and answer selection built from online health forums. One of the oldest entries in th

DataDEL

Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished

HELM

The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedOdyssey

Long-context medical evaluation running up to 200,000 tokens.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities