Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account.
Wu, Wu and Zou, 20261 verified conflict
Key facts
| What it measures | Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account. |
| Who should care | Anyone who wants verified clinician judgment on real questions rather than exam scores. |
| Use it when | You want to know what doctors actually prefer on the questions they really ask. |
| Do not use it for | You need correctness. Preference is not accuracy, and clinicians cite depth and clarity more often than factual accuracy when explaining their choices. |
| Score scale | Bradley-Terry and Elo rating with bootstrapped 95% confidence intervals and pairwise p-values. Rank gaps inside the interval are noise. |
| Grader method | Bradley-Terry and Elo with bootstrapped 95% CI |
| First released | 2025 |
| Languages | English |
| Use cases | clinical decision support, patient communication |
| Citation | Wu, Wu and Zou, 2026 |
| Status | active |
Who is behind MedArena
Verified conflicts of interest
Independence assessment
Run at Stanford University by Eric Wu, Kevin Wu and James Zou in the Zou Lab. No commercial steward identified. Rankings are recomputed at intervals that depend on preference volume rather than on a fixed published cadence.
Scope
1,571 clinician preferences across 12 models collected to 1 November 2025. Roughly a third of submitted questions are factual recall; the majority are treatment selection, documentation or patient communication, and about a fifth are multi-turn.
Frequently asked questions
What does MedArena measure?
Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access is restricted to professionals with an NPI or Doximity account.
Who should use MedArena?
Anyone who wants verified clinician judgment on real questions rather than exam scores. You want to know what doctors actually prefer on the questions they really ask.
What should MedArena not be used for?
You need correctness. Preference is not accuracy, and clinicians cite depth and clarity more often than factual accuracy when explaining their choices.
How are MedArena scores reported?
Bradley-Terry and Elo rating with bootstrapped 95% confidence intervals and pairwise p-values. Rank gaps inside the interval are noise. Scores from different benchmarks are not comparable to each other.
Is MedArena independent?
Not fully. Update cadence is volume-driven, not a fixed schedule.
Compare with
Simulated clinical encounters where the model must gather information over several turns rather than answer on
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
