Quantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer.

MedR-Bench is a benchmark in healthcare AI. Quantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer. Scores are reported as: Reasoning-step accuracy and completeness.

Audited by MedCheckQiu et al., 2025

Key facts

What it measuresQuantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer.
Who should careTeams that need to audit why a model reached a conclusion.
Use it whenReasoning transparency evaluation.
Do not use it forFast throughput testing.
Score scaleReasoning-step accuracy and completeness.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesclinical decision support, safety
CitationQiu et al., 2025
Statusactive

Paper

Frequently asked questions

What does MedR-Bench measure?

Quantifies reasoning quality on real-world clinical cases, scoring the path rather than only the answer.

Who should use MedR-Bench?

Teams that need to audit why a model reached a conclusion. Reasoning transparency evaluation.

What should MedR-Bench not be used for?

Fast throughput testing.

How are MedR-Bench scores reported?

Reasoning-step accuracy and completeness. Scores from different benchmarks are not comparable to each other.

Is MedR-Bench independent?

No conflicts of interest have been verified for MedR-Bench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities