Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.

HealthBench is a benchmark in healthcare AI. Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones. Scores are reported as: Rubric score, percent of criteria met.

Audited by MedCheckArora et al., 20251 verified conflict

Key facts

What it measuresPhysician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.
Who should careAnyone evaluating a general model for health questions from patients or clinicians.
Use it whenYou want physician-designed rubric grading rather than multiple choice.
Do not use it forYou need a vendor-neutral referee. The benchmark author is also a model vendor.
Score scaleRubric score, percent of criteria met.
Grader methodunknown
First released2025
LanguagesEnglish
Use casespatient communication, clinical decision support, safety
CitationArora et al., 2025
Statusactive

Paper

Who is behind HealthBench

creator and commercial
OpenAI
Sells: Frontier model provider
Benchmark creator defines the tasks and the scoring system and also reports its own product results.
Source, verified

Verified conflicts of interest

Created by a frontier model vendor that also reports its own results on it

Frequently asked questions

What does HealthBench measure?

Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.

Who should use HealthBench?

Anyone evaluating a general model for health questions from patients or clinicians. You want physician-designed rubric grading rather than multiple choice.

What should HealthBench not be used for?

You need a vendor-neutral referee. The benchmark author is also a model vendor.

How are HealthBench scores reported?

Rubric score, percent of criteria met. Scores from different benchmarks are not comparable to each other.

Is HealthBench independent?

Not fully. Created by a frontier model vendor that also reports its own results on it.

Compare with

CHBench

Chinese-language evaluation of physical and mental health knowledge in large models.

LMArena Medicine and Healthcare

Open public voting on model answers, filtered to a medicine and healthcare occupational category.

MedArena

Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access

MedExQA

Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedJourney

Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isol

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities