Open public voting on model answers, filtered to a medicine and healthcare occupational category.

LMArena Medicine and Healthcare is a arena in healthcare AI. Open public voting on model answers, filtered to a medicine and healthcare occupational category. Scores are reported as: Elo with published 95% confidence intervals.

LMArena, 20262 verified conflicts

Key facts

What it measuresOpen public voting on model answers, filtered to a medicine and healthcare occupational category.
Who should carePeople tracking general preference trends, not clinical buyers.
Use it whenYou want a fast read on which models people prefer on health prompts.
Do not use it forYou are making a clinical procurement decision. Voters are not verified clinicians.
Score scaleElo with published 95% confidence intervals.
Grader methodPairwise human preference Elo with published 95% CI
First released2026
LanguagesEnglish
Use casesgeneral, patient communication
CitationLMArena, 2026
Statusactive

Official site · Leaderboard

Verified conflicts of interest

Open voting is optimizable by vendors
Voters are not verified clinicians

Independence assessment

Open public voting. Known field-wide concern that providers can optimize for preference signal; Style Control is the partial mitigation. Not clinician-restricted, so preference does not equal clinical correctness.

Scope

Occupational category leaderboard within LMArena covering medicine and healthcare prompts. Measures human preference, not clinical accuracy.

Frequently asked questions

What does LMArena Medicine and Healthcare measure?

Open public voting on model answers, filtered to a medicine and healthcare occupational category.

Who should use LMArena Medicine and Healthcare?

People tracking general preference trends, not clinical buyers. You want a fast read on which models people prefer on health prompts.

What should LMArena Medicine and Healthcare not be used for?

You are making a clinical procurement decision. Voters are not verified clinicians.

How are LMArena Medicine and Healthcare scores reported?

Elo with published 95% confidence intervals. Scores from different benchmarks are not comparable to each other.

Is LMArena Medicine and Healthcare independent?

Not fully. Open voting is optimizable by vendors; Voters are not verified clinicians.

Compare with

CHBench

Chinese-language evaluation of physical and mental health knowledge in large models.

HealthBench

Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing

MedArena

Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access

MedExQA

Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedJourney

Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isol

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities