A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
Audited by MedCheckFan et al., COLING 2025
Key facts
| What it measures | A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction. |
| Who should care | Teams building conversational clinical agents. |
| Use it when | Testing behavior in a full simulated interaction. |
| Do not use it for | Static evaluation. |
| Score scale | Multi-view scoring across simulated roles. |
| Grader method | unknown |
| First released | 2025 |
| Languages | Chinese, English |
| Use cases | agentic, patient communication |
| Citation | Fan et al., COLING 2025 |
| Status | active |
Frequently asked questions
What does MVME (AI Hospital) measure?
A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.
Who should use MVME (AI Hospital)?
Teams building conversational clinical agents. Testing behavior in a full simulated interaction.
What should MVME (AI Hospital) not be used for?
Static evaluation.
How are MVME (AI Hospital) scores reported?
Multi-view scoring across simulated roles. Scores from different benchmarks are not comparable to each other.
Is MVME (AI Hospital) independent?
No conflicts of interest have been verified for MVME (AI Hospital) in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Clinical reasoning in the emergency room, built from real de-identified ER records.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa
Interactive sequential benchmarking that follows a case through stages the way a real clinic does.
Whether a model asks the right follow-up question when it does not have enough information, instead of guessin
