Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.
Audited by MedCheckJin et al., EMNLP-IJCNLP 2019
Key facts
| What it measures | Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts. |
| Who should care | Teams building literature research tools. |
| Use it when | Evaluating literature comprehension. |
| Do not use it for | Evaluating clinical reasoning. |
| Score scale | Percent accuracy across three answer classes. |
| Grader method | unknown |
| First released | 2019 |
| Languages | English |
| Use cases | research |
| Citation | Jin et al., EMNLP-IJCNLP 2019 |
| Status | active |
Frequently asked questions
What does PubMedQA measure?
Yes, no and maybe answers to biomedical research questions drawn from PubMed abstracts.
Who should use PubMedQA?
Teams building literature research tools. Evaluating literature comprehension.
What should PubMedQA not be used for?
Evaluating clinical reasoning.
How are PubMedQA scores reported?
Percent accuracy across three answer classes. Scores from different benchmarks are not comparable to each other.
Is PubMedQA independent?
No conflicts of interest have been verified for PubMedQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Chinese medical question and answer selection built from online health forums. One of the oldest entries in th
Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished
The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Long-context medical evaluation running up to 200,000 tokens.
