Healthcare AI benchmark ownership and governance | Healthcare Discovery
|

What Is HealthBench? How OpenAI’s Medical AI Benchmark Tests Models Against Doctors

The medical question arrives without ceremony. A patient has symptoms, a clinician has an incomplete record, or a researcher needs to reconcile evidence that does not quite agree. The answer must do more than retrieve a fact. It has to notice what is missing, distinguish danger from inconvenience, communicate uncertainty, and know when the safest recommendation is to stop typing and send someone for care.

Presented By Our Partners

That is the problem HealthBench was designed to measure.

HealthBench is a family of medical-AI evaluations created by OpenAI with physicians. Instead of asking a model to choose A, B, C, or D on a medical examination, it presents realistic, open-ended health conversations. Physicians write a customized scoring rubric for each conversation, identifying what a strong answer should include, what it should avoid, and which omissions could be dangerous.

The short answer is that HealthBench measures the quality and safety of written AI responses in defined health scenarios. It does not determine whether an AI system can replace a physician. A high score can show that a model handled a set of conversations well under specified conditions. It cannot show that the model can examine a patient, establish a diagnosis in the real world, manage treatment over time, or accept responsibility when something goes wrong.

That distinction became more important in 2026, when HealthBench Professional reported that an AI system produced higher-scoring written responses than a strong physician-written baseline. The result is consequential. It is also much narrower than “AI beat doctors.”

Why medical AI needed a different kind of test

For years, medical language models were commonly evaluated using exam questions. Those tests were useful: they showed whether a model could retrieve biomedical knowledge and reason toward a predetermined answer. But a patient rarely enters a clinic carrying four possible diagnoses and asking the doctor to select one.

Real health conversations are untidy. People omit medication names, confuse timelines, minimize dangerous symptoms, or ask for certainty where certainty is impossible. Clinicians must decide what to ask next, what not to assume, how much detail is useful, and when escalation matters more than explanation.

The original HealthBench, released in 2025, tried to capture that complexity. It contained 5,000 multi-turn conversations between models and either individual users or healthcare professionals. A group of 262 physicians across 26 specialties and 60 countries contributed to the work. Together, they produced 48,562 conversation-specific rubric criteria.

Those criteria make HealthBench different from a conventional test. An answer might earn credit for recognizing a red flag, asking for missing context, explaining uncertainty, or adapting advice to the user’s location. It might lose credit for confidently inventing a diagnosis, omitting an urgent warning, or burying the useful answer beneath unnecessary technical language.

The benchmark therefore asks a more realistic question than “Does the model know the fact?” It asks: “Did the model behave helpfully and safely in this particular conversation?”

The four parts of the HealthBench family

HealthBench is now better understood as a family of related evaluations.

HealthBench is the broad foundation: 5,000 realistic health conversations scored against physician-written criteria.

HealthBench Hard isolates especially difficult cases on which frontier models have substantial room to improve.

HealthBench Consensus emphasizes criteria that multiple physicians validate, allowing researchers to compare model-based grading with physician judgment.

HealthBench Professional, introduced in 2026, narrows the lens to work clinicians might bring to an AI system. Its tasks cover three categories: care consultation; writing and documentation; and medical research.

The distinction matters. A consumer asking about chest pain and a hematologist asking for help synthesizing a difficult literature question are both “health” interactions, but they test different capabilities and create different risks.

How HealthBench Professional was built

HealthBench Professional began with 15,079 candidate examples created through clinician use and deliberate stress testing. The final benchmark contains 525 physician-authored tasks selected for quality, representativeness, and difficulty.

The researchers did not try to reproduce the natural frequency of easy and hard clinical questions. Difficult examples were enriched by roughly 3.5 times compared with the candidate pool, and about one-third of the final tasks came from deliberate physician red teaming. That makes the benchmark useful for finding failure modes, but it also means its average score should not be interpreted as the percentage of ordinary clinical work a model can perform correctly.

Each task contains a conversation followed by criteria written for that specific situation. Criteria can carry positive or negative point values. A model-based grader determines whether the response satisfies each criterion, and the points are combined into a score.

This design offers granularity that multiple-choice tests cannot. It also introduces judgment calls. The physicians decide what belongs in the rubric. A separate AI model decides whether each criterion was met. The benchmark is therefore not a direct measurement device like a thermometer. It is a structured system of expert judgments, encoded into rubrics and applied at scale by another model.

Featured Partner

Invest in the Infrastructure Behind Modern Medicine

As healthcare expands beyond hospital walls, the buildings and campuses supporting that shift are generating compelling returns for investors who move early. The Healthcare Real Estate Fund offers qualified investors direct access to a curated portfolio of medical office, outpatient, and specialty care facilities.

Learn More →

Why answer length became a scientific problem

Open-ended benchmarks have an awkward vulnerability: longer answers often mention more rubric items simply because they contain more words. A model can sometimes raise its score by becoming verbose without becoming more clinically useful.

The HealthBench Professional researchers tested models at different verbosity settings and found that greater length could increase the unadjusted score. They therefore report a length-adjusted result. Around the common response range, the published method applies a penalty of 1.47 points for every additional 500 characters above 2,000, with corresponding adjustments for shorter answers.

That correction is scientifically sensible, but it also complicates public leaderboards. A score can depend not only on the underlying model, but on its prompt, reasoning setting, product interface, browsing tools, answer length, and the grader used. Two rows carrying similar model names may not represent the same system.

This is why a trustworthy medical-AI ranking cannot simply copy the largest number from a chart. It must record the model version, test date, reasoning effort, tools, product harness, grader, length adjustment, and source.

Did AI outperform doctors?

On HealthBench Professional, the original 2026 paper reports that GPT-5.4 operating inside ChatGPT for Clinicians scored 59.0. Physician-written responses scored 43.7. The difference was statistically significant, with a reported p value of 3.7 × 10^-10.

The physician comparison was not casual. Responses were written for every benchmark task by specialty-matched physicians who had unlimited time and internet access, but who were not allowed to use AI. The researchers intended this to be a strong human baseline.

That makes the finding worth taking seriously. Under the benchmark’s rules, the productized AI system produced responses that satisfied more of the physician-written rubric criteria than the physician-written responses did.

But the experiment did not place an AI system and a doctor in parallel clinics and compare patient outcomes. It evaluated the next written response in a conversation. The AI did not perform a physical examination, observe a patient’s condition changing, coordinate care, obtain informed consent, or live with the consequences of a decision.

There is another important detail: the highest score came from GPT-5.4 inside ChatGPT for Clinicians, not from the base GPT-5.4 model alone. The product system scored 59.0, while base GPT-5.4 scored 48.1 and a simpler GPT-5.4 browsing configuration scored 45.8. The surrounding product—its instructions, retrieval, tools, and workflow—was part of the result.

The clean conclusion is therefore not “AI is a better doctor.” It is this: on a difficult set of clinician-facing written tasks, a specialized AI product outscored a strong physician-written baseline according to the benchmark’s rubric-based method.

Which model leads HealthBench Professional in 2026?

There is not yet one clean, independently maintained leaderboard that makes every 2026 result directly comparable.

In the original HealthBench Professional paper, GPT-5.4 inside ChatGPT for Clinicians led the evaluated field at 59.0. Base GPT-5.4 scored 48.1, a simple GPT-5.4 browsing harness scored 45.8, and the specialty-matched physician baseline scored 43.7.

Meta later published a different evaluation snapshot in its Muse Spark 1.1 report. Meta listed Muse Spark 1.1 at 59.3, Claude Opus 4.8 at 55.8, the original Muse Spark at 54.1, GPT-5.5 at 51.8, and Gemini 3.1 Pro at 41.6.

Those numbers should not be collapsed into a single definitive ranking. The snapshots were produced by different model developers, at different times, with different comparison sets and configurations. Meta states that Muse Spark 1.1 was run through the Meta Model API at xhigh reasoning effort, that GPT-5.5 used xhigh effort, Claude used max effort, and Gemini used high effort. It also states that its HealthBench Professional results used GPT-5.4 at low reasoning effort as grader and the paper’s length-normalized scoring approach.

For now, the defensible verdict is two-part: ChatGPT for Clinicians led the original published study, while Muse Spark 1.1 posted the highest score in Meta’s later developer-produced snapshot. Neither result has yet earned the status of a neutral, universal championship.

The benchmark-maxing problem

Once a benchmark becomes important, developers have incentives to optimize for it. That does not automatically invalidate the result; medical education itself involves practicing against known standards. But optimization can weaken a benchmark’s ability to predict performance outside the test.

HealthBench Professional is open, and its examples were selected partly for difficulty against contemporary OpenAI models. Model developers can study its structure. Product teams can tune prompts, verbosity, browsing behavior, or answer format to satisfy more rubric criteria. A product may become genuinely better—or merely better adapted to the scoreboard.

Developer-produced evaluations add another layer of caution. OpenAI created HealthBench and reports results for OpenAI systems. Meta reports results for Muse Spark 1.1. Those reports are valuable primary sources for methods and claimed performance, but they are not independent confirmation.

A credible ranking should therefore separate at least four questions:

  1. Who created the benchmark?
  2. Who evaluated the model?
  3. Was the result independently reproduced?
  4. Does performance persist across different medical tasks, graders, prompts, and datasets?

The leaderboard is the beginning of the investigation, not the verdict.

What HealthBench scores do—and do not—tell us

A HealthBench result can provide evidence about written response quality under defined conditions. Depending on the variant, it may test accuracy, completeness, communication, context awareness, escalation, documentation, research synthesis, or safety-related behavior.

It does not automatically establish:

  • diagnostic accuracy in real patients;
  • improved morbidity, mortality, or quality of life;
  • safe autonomous treatment decisions;
  • performance across every specialty, language, or health system;
  • regulatory authorization as a medical device;
  • freedom from hallucinations or dangerous omissions;
  • the ability to replace a clinician’s examination, relationship, judgment, or accountability.

The most useful interpretation is task-specific. A model may be excellent at drafting a structured note, strong at summarizing research, uneven at triage, and unsafe when important context is missing. There may never be one universally “best AI doctor.” The more defensible question is which system performs best for a defined purpose, under what conditions, and according to whose evidence.

Frequently asked questions

What is HealthBench?

HealthBench is a physician-designed benchmark family for evaluating the quality and safety of AI responses in realistic health conversations. Responses are scored against criteria written for each individual conversation.

Is HealthBench a medical licensing examination?

No. It uses open-ended conversations rather than conventional multiple-choice licensing questions, and it evaluates model behavior as well as factual content.

What is HealthBench Professional?

HealthBench Professional is a 2026 benchmark containing 525 physician-authored clinician tasks across care consultation, writing and documentation, and medical research.

Did ChatGPT beat doctors on HealthBench Professional?

GPT-5.4 inside ChatGPT for Clinicians scored higher than physician-written responses under the benchmark’s rubric-based method. That supports superiority on the tested written tasks, not superiority as a doctor in general clinical practice.

Were the physicians allowed to use the internet?

Yes. The specialty-matched physicians had unbounded time and web access but were instructed not to use AI.

Does a score of 59 mean 59% diagnostic accuracy?

No. It is a rubric-derived benchmark score and should not be interpreted as the percentage of diagnoses made correctly.

Can models be trained to perform well on HealthBench?

Yes. Because the benchmark is influential and its structure is public, benchmark-specific optimization and contamination are legitimate concerns. Independent and private evaluations help test whether gains generalize.

Can HealthBench identify the best AI doctor?

It can contribute evidence about particular written tasks. A responsible ranking must also consider safety, calibration, deployment context, independent validation, cost, tools, and real clinical outcomes.

Before acting on an AI health answer

An AI response is one opinion, not a complete clinical evaluation. Even an excellent benchmark score cannot give a model the ability to examine you, know every relevant detail, or assume responsibility for your care.

A second opinion is especially valuable from a qualified human clinician. Bring the AI response, the sources it cited, and the questions it raised. A local doctor can place that information within your medical history, symptoms, examination, and treatment options.

If symptoms may be urgent or life-threatening, seek emergency or urgent medical care rather than waiting for an AI conversation or routine appointment.

Sources

This article is educational and does not provide diagnosis or medical advice. Benchmark results describe performance under specific test conditions and may change as models, products, and evaluation methods evolve.

Free Daily Briefing

The Latest Longevity Science.
Delivered Every Morning.

Join researchers, physicians, and health professionals getting daily breakthroughs in AI-driven medicine, epigenetics, and longevity research.

Support the research that powers this editorial

No spam. Unsubscribe anytime. We respect your inbox.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *