A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions.

MedAgentBench is a benchmark in healthcare AI. A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions. Scores are reported as: Task completion rate in the simulated EHR.

Audited by MedCheckJiang et al., 2025

Key facts

What it measuresA realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions.
Who should careTeams deploying agents that will touch an EHR.
Use it whenEvaluating tool use and action safety inside a record system.
Do not use it forEvaluating knowledge or conversation.
Score scaleTask completion rate in the simulated EHR.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesagentic, documentation
CitationJiang et al., 2025
Statusactive

Paper

Frequently asked questions

What does MedAgentBench measure?

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions.

Who should use MedAgentBench?

Teams deploying agents that will touch an EHR. Evaluating tool use and action safety inside a record system.

What should MedAgentBench not be used for?

Evaluating knowledge or conversation.

How are MedAgentBench scores reported?

Task completion rate in the simulated EHR. Scores from different benchmarks are not comparable to each other.

Is MedAgentBench independent?

No conflicts of interest have been verified for MedAgentBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

MedAgentsBench

Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa

MedChain

Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

MediQ

Whether a model asks the right follow-up question when it does not have enough information, instead of guessin

MVME (AI Hospital)

A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities