Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account.

MedArena is a arena in healthcare AI. Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account. Scores are reported as: Bradley-Terry and Elo rating with bootstrapped 95% confidence intervals and pairwise p-values. Rank gaps inside the interval are noise.

Wu, Wu and Zou, 20261 verified conflict

Key facts

What it measuresPracticing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account.
Who should careAnyone who wants verified clinician judgment on real questions rather than exam scores.
Use it whenYou want to know what doctors actually prefer on the questions they really ask.
Do not use it forYou need correctness. Preference is not accuracy, and clinicians cite depth and clarity more often than factual accuracy when explaining their choices.
Score scaleBradley-Terry and Elo rating with bootstrapped 95% confidence intervals and pairwise p-values. Rank gaps inside the interval are noise.
Grader methodBradley-Terry and Elo with bootstrapped 95% CI
First released2025
LanguagesEnglish
Use casesclinical decision support, patient communication
CitationWu, Wu and Zou, 2026
Statusactive

Official site · Paper

Who is behind MedArena

creator
Stanford HAI
Source, verified
maintainer
Stanford University Zou Lab
Source, verified

Verified conflicts of interest

Update cadence is volume-driven, not a fixed schedule

Independence assessment

Run at Stanford University by Eric Wu, Kevin Wu and James Zou in the Zou Lab. No commercial steward identified. Rankings are recomputed at intervals that depend on preference volume rather than on a fixed published cadence.

Scope

1,571 clinician preferences across 12 models collected to 1 November 2025. Roughly a third of submitted questions are factual recall; the majority are treatment selection, documentation or patient communication, and about a fifth are multi-turn.

Frequently asked questions

What does MedArena measure?

Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account.

Who should use MedArena?

Anyone who wants verified clinician judgment on real questions rather than exam scores. You want to know what doctors actually prefer on the questions they really ask.

What should MedArena not be used for?

You need correctness. Preference is not accuracy, and clinicians cite depth and clarity more often than factual accuracy when explaining their choices.

How are MedArena scores reported?

Bradley-Terry and Elo rating with bootstrapped 95% confidence intervals and pairwise p-values. Rank gaps inside the interval are noise. Scores from different benchmarks are not comparable to each other.

Is MedArena independent?

Not fully. Update cadence is volume-driven, not a fixed schedule.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities