Whether a model asks the right follow-up question when it does not have enough information, instead of guessing.
Audited by MedCheckLi et al., NeurIPS 2024
Key facts
| What it measures | Whether a model asks the right follow-up question when it does not have enough information, instead of guessing. |
| Who should care | Anyone building triage or intake tools. |
| Use it when | Testing information-seeking behavior, which most benchmarks never measure. |
| Do not use it for | Static question answering. |
| Score scale | Question quality and downstream diagnostic accuracy. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | agentic, safety |
| Citation | Li et al., NeurIPS 2024 |
| Status | active |
Frequently asked questions
What does MediQ measure?
Whether a model asks the right follow-up question when it does not have enough information, instead of guessing.
Who should use MediQ?
Anyone building triage or intake tools. Testing information-seeking behavior, which most benchmarks never measure.
What should MediQ not be used for?
Static question answering.
How are MediQ scores reported?
Question quality and downstream diagnostic accuracy. Scores from different benchmarks are not comparable to each other.
Is MediQ independent?
No conflicts of interest have been verified for MediQ in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Clinical reasoning in the emergency room, built from real de-identified ER records.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa
Interactive sequential benchmarking that follows a case through stages the way a real clinic does.
A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
