Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

MedChain is a benchmark in healthcare AI. Interactive sequential benchmarking that follows a case through stages the way a real clinic does. Scores are reported as: Stage-level and end-to-end accuracy.

Audited by MedCheckLiu et al., 2024

Key facts

What it measuresInteractive sequential benchmarking that follows a case through stages the way a real clinic does.
Who should careTeams building multi-turn clinical tools.
Use it whenTesting whether performance holds across a whole encounter, not one question.
Do not use it forSingle-turn comparison.
Score scaleStage-level and end-to-end accuracy.
Grader methodunknown
First released2024
LanguagesChinese, English
Use casesagentic, clinical decision support
CitationLiu et al., 2024
Statusactive

Paper

Frequently asked questions

What does MedChain measure?

Interactive sequential benchmarking that follows a case through stages the way a real clinic does.

Who should use MedChain?

Teams building multi-turn clinical tools. Testing whether performance holds across a whole encounter, not one question.

What should MedChain not be used for?

Single-turn comparison.

How are MedChain scores reported?

Stage-level and end-to-end accuracy. Scores from different benchmarks are not comparable to each other.

Is MedChain independent?

No conflicts of interest have been verified for MedChain in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

ER-REASON

Clinical reasoning in the emergency room, built from real de-identified ER records.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedAgentsBench

Compares thinking models and multi-agent frameworks on complex medical reasoning that single-pass answering fa

MediQ

Whether a model asks the right follow-up question when it does not have enough information, instead of guessin

MVME (AI Hospital)

A multi-agent simulator where models play clinician, patient and examiner roles through a full interaction.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities