A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions.
Audited by MedCheckJiang et al., 2025
Key facts
| What it measures | A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions. |
| Who should care | Teams deploying agents that will touch an EHR. |
| Use it when | Evaluating tool use and action safety inside a record system. |
| Do not use it for | Evaluating knowledge or conversation. |
| Score scale | Task completion rate in the simulated EHR. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | agentic, documentation |
| Citation | Jiang et al., 2025 |
| Status | active |
Frequently asked questions
What does MedAgentBench measure?
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, not answering questions.
Who should use MedAgentBench?
Teams deploying agents that will touch an EHR. Evaluating tool use and action safety inside a record system.
What should MedAgentBench not be used for?
Evaluating knowledge or conversation.
How are MedAgentBench scores reported?
Task completion rate in the simulated EHR. Scores from different benchmarks are not comparable to each other.
Is MedAgentBench independent?
No conflicts of interest have been verified for MedAgentBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Clinical reasoning in the emergency room, built from real de-identified ER records.
Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa
Interactive sequential benchmarking that follows a case through stages the way a real clinic does.
Whether a model asks the right follow-up question when it does not have enough information, instead of guessin
A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
