Whether a model asks the right follow-up question when it does not have enough information, instead of guessing.

MediQ is a benchmark in healthcare AI. Whether a model asks the right follow-up question when it does not have enough information, instead of guessing. Scores are reported as: Question quality and downstream diagnostic accuracy.

Audited by MedCheckLi et al., NeurIPS 2024

Key facts

What it measuresWhether a model asks the right follow-up question when it does not have enough information, instead of guessing.
Who should careAnyone building triage or intake tools.
Use it whenTesting information-seeking behavior, which most benchmarks never measure.
Do not use it forStatic question answering.
Score scaleQuestion quality and downstream diagnostic accuracy.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesagentic, safety
CitationLi et al., NeurIPS 2024
Statusactive

Frequently asked questions

What does MediQ measure?

Whether a model asks the right follow-up question when it does not have enough information, instead of guessing.

Who should use MediQ?

Anyone building triage or intake tools. Testing information-seeking behavior, which most benchmarks never measure.

What should MediQ not be used for?

Static question answering.

How are MediQ scores reported?

Question quality and downstream diagnostic accuracy. Scores from different benchmarks are not comparable to each other.

Is MediQ independent?

No conflicts of interest have been verified for MediQ in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedAgentsBench

Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa

MedChain

Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

MVME (AI Hospital)

A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities