Simulated clinical encounters where the model must gather information over several turns rather than answer one question.

AgentClinic is a benchmark in healthcare AI. Simulated clinical encounters where the model must gather information over several turns rather than answer one question. Scores are reported as: Task success rate across simulated encounters.

Audited by MedCheckSchmidgall et al., 2024

Key facts

What it measuresSimulated clinical encounters where the model must gather information over several turns rather than answer one question.
Who should careTeams building agents rather than chatbots.
Use it whenEvaluating sequential decision making and information gathering.
Do not use it forYou want a fast single-number comparison.
Score scaleTask success rate across simulated encounters.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesagentic, clinical decision support
CitationSchmidgall et al., 2024
Statusactive

Paper

Frequently asked questions

What does AgentClinic measure?

Simulated clinical encounters where the model must gather information over several turns rather than answer one question.

Who should use AgentClinic?

Teams building agents rather than chatbots. Evaluating sequential decision making and information gathering.

What should AgentClinic not be used for?

You want a fast single-number comparison.

How are AgentClinic scores reported?

Task success rate across simulated encounters. Scores from different benchmarks are not comparable to each other.

Is AgentClinic independent?

No conflicts of interest have been verified for AgentClinic in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedAgentsBench

Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa

MedChain

Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

MediQ

Whether a model asks the right follow-up question when it does not have enough information, instead of guessin

MVME (AI Hospital)

A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities