Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.
Audited by MedCheckArora et al., 20251 verified conflict
Key facts
| What it measures | Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones. |
| Who should care | Anyone evaluating a general model for health questions from patients or clinicians. |
| Use it when | You want physician-designed rubric grading rather than multiple choice. |
| Do not use it for | You need a vendor-neutral referee. The benchmark author is also a model vendor. |
| Score scale | Rubric score, percent of criteria met. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | patient communication, clinical decision support, safety |
| Citation | Arora et al., 2025 |
| Status | active |
Who is behind HealthBench
Verified conflicts of interest
Frequently asked questions
What does HealthBench measure?
Physician-written rubrics score open-ended answers to realistic health conversations, including patient-facing ones.
Who should use HealthBench?
Anyone evaluating a general model for health questions from patients or clinicians. You want physician-designed rubric grading rather than multiple choice.
What should HealthBench not be used for?
You need a vendor-neutral referee. The benchmark author is also a model vendor.
How are HealthBench scores reported?
Rubric score, percent of criteria met. Scores from different benchmarks are not comparable to each other.
Is HealthBench independent?
Not fully. Created by a frontier model vendor that also reports its own results on it.
Compare with
Chinese-language evaluation of physical and mental health knowledge in large models.
Open public voting on model answers, filtered to a medicine and healthcare occupational category.
Practicing clinicians submit their own real questions and pick which of two model answers they prefer. Access
Medical questions paired with multiple valid expert explanations, so a model is judged on reasoning as well as
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isol
