Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fails.
Audited by MedCheckTang et al., 2025
Key facts
| What it measures | Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fails. |
| Who should care | Teams choosing between reasoning models and agent scaffolds. |
| Use it when | Deciding whether an agent framework earns its cost over a single model call. |
| Do not use it for | Simple factual question answering. |
| Score scale | Accuracy on hard multi-step cases, reported by framework. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | agentic, clinical decision support |
| Citation | Tang et al., 2025 |
| Status | active |
Frequently asked questions
What does MedAgentsBench measure?
Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fails.
Who should use MedAgentsBench?
Teams choosing between reasoning models and agent scaffolds. Deciding whether an agent framework earns its cost over a single model call.
What should MedAgentsBench not be used for?
Simple factual question answering.
How are MedAgentsBench scores reported?
Accuracy on hard multi-step cases, reported by framework. Scores from different benchmarks are not comparable to each other.
Is MedAgentsBench independent?
No conflicts of interest have been verified for MedAgentsBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Clinical reasoning in the emergency room, built from real de-identified ER records.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Interactive sequential benchmarking that follows a case through stages the way a real clinic does.
Whether a model asks the right follow-up question when it does not have enough information, instead of guessin
A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
