A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

MVME (AI Hospital) is a benchmark in healthcare AI. A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction. Scores are reported as: Multi-view scoring across simulated roles.

Audited by MedCheckFan et al., COLING 2025

Key facts

What it measuresA multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
Who should careTeams building conversational clinical agents.
Use it whenTesting behavior in a full simulated interaction.
Do not use it forStatic evaluation.
Score scaleMulti-view scoring across simulated roles.
Grader methodunknown
First released2025
LanguagesChinese, English
Use casesagentic, patient communication
CitationFan et al., COLING 2025
Statusactive

Frequently asked questions

What does MVME (AI Hospital) measure?

A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

Who should use MVME (AI Hospital)?

Teams building conversational clinical agents. Testing behavior in a full simulated interaction.

What should MVME (AI Hospital) not be used for?

Static evaluation.

How are MVME (AI Hospital) scores reported?

Multi-view scoring across simulated roles. Scores from different benchmarks are not comparable to each other.

Is MVME (AI Hospital) independent?

No conflicts of interest have been verified for MVME (AI Hospital) in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedAgentsBench

Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa

MedChain

Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

MediQ

Whether a model asks the right follow-up question when it does not have enough information, instead of guessin

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities