Open public voting on model answers, filtered to a medicine and healthcare occupational category.
LMArena, 20262 verified conflicts
Key facts
| What it measures | Open public voting on model answers, filtered to a medicine and healthcare occupational category. |
| Who should care | People tracking general preference trends, not clinical buyers. |
| Use it when | You want a fast read on which models people prefer on health prompts. |
| Do not use it for | You are making a clinical procurement decision. Voters are not verified clinicians. |
| Score scale | Elo with published 95% confidence intervals. |
| Grader method | Pairwise human preference Elo with published 95% CI |
| First released | 2026 |
| Languages | English |
| Use cases | general, patient communication |
| Citation | LMArena, 2026 |
| Status | active |
Verified conflicts of interest
Independence assessment
Open public voting. Known field-wide concern that providers can optimize for preference signal; Style Control is the partial mitigation. Not clinician-restricted, so preference does not equal clinical correctness.
Scope
Occupational category leaderboard within LMArena covering medicine and healthcare prompts. Measures human preference, not clinical accuracy.
Frequently asked questions
What does LMArena Medicine and Healthcare measure?
Open public voting on model answers, filtered to a medicine and healthcare occupational category.
Who should use LMArena Medicine and Healthcare?
People tracking general preference trends, not clinical buyers. You want a fast read on which models people prefer on health prompts.
What should LMArena Medicine and Healthcare not be used for?
You are making a clinical procurement decision. Voters are not verified clinicians.
How are LMArena Medicine and Healthcare scores reported?
Elo with published 95% confidence intervals. Scores from different benchmarks are not comparable to each other.
Is LMArena Medicine and Healthcare independent?
Not fully. Open voting is optimizable by vendors; Voters are not verified clinicians.
Compare with
Chinese-language evaluation of physical and mental health knowledge in large models.
Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing
Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access
Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isol
