Healthcare AI benchmark registry: who owns the scoreboards that grade medical AI | Healthcare Discovery
| |

Who Decides Which AI Is Good Enough for Medicine?

Sixty-three separate scoreboards grade medical AI. Several are built or maintained by the companies whose models they rank. We mapped all of them.

Presented By Our Partners

No single body decides which artificial intelligence is good enough to use in medicine. Sixty-three separate instruments score medical AI, and several of them are created or maintained by companies with a commercial stake in the results. Scores from different instruments cannot be compared to one another. The most widely cited clinical AI leaderboard has not been updated in nearly three months.

That is the short answer. The longer one is stranger, and it starts with a question almost nobody asks out loud.

Who decides which AI is good enough for medicine?

Ask who approves a drug and the answer is immediate. The Food and Drug Administration. Ask who approves a medical device and the answer is the same agency through a different pathway. Ask who decides whether a language model is good enough to summarize a discharge summary, draft a patient message, or help a physician think through a differential, and the answer dissolves.

There is no regulator for this. What exists instead is a loose federation of benchmarks: standardized sets of medical tasks that researchers use to score models. A model that does well becomes the one health systems buy. A benchmark that becomes influential shapes what model developers optimize for. Nobody elected any of them.

We spent the last several weeks building a structured registry of every instrument we could identify, starting from a peer-reviewed audit of the field published at ACL 2026 and adding the living leaderboards, clinician preference arenas, aggregators and certification bodies that the audit excluded. Sixty-three entities. For each one we recorded what it measures, who created it, who funds it, who maintains the code, whether that maintainer sells a product that depends on the benchmark, what units the scores use, and how many days old the published results are.

Three findings surprised us.

Why do two medical AI leaderboards rank the same models differently?

Because they are measuring different things in different units, and almost nothing in how the results are presented tells you that.

MedHELM, the most substantial clinical task benchmark, reports a mean win rate between zero and one. Its current leader sits at 0.652. A commercial aggregator publishing a healthcare model ranking reports its leader at 45.1 on an undocumented composite scale. MedArena, a Stanford platform where verified clinicians choose between two model answers, reports Elo ratings with confidence intervals, where a gap of a few points is statistical noise rather than a difference.

None of these three numbers can be compared to either of the others. They are not the same measurement expressed differently. They are different measurements entirely, in the way that a blood pressure reading and a body temperature are both numbers about a patient and tell you nothing about each other.

This is the single most common error in healthcare AI procurement, and it is not the buyer’s fault. Leaderboards present ranked tables with scores in a column, a visual grammar that invites comparison. Almost none of them state their units in language a non-specialist would notice.

Who owns the healthcare AI scoreboards?

This is the layer nobody publishes, and it is where the registry found the most.

HealthBench, one of the most influential medical AI evaluations, was created by OpenAI working with physicians. The design is serious: 5,000 realistic multi-turn health conversations, 262 physicians across 26 specialties and 60 countries, and 48,562 conversation-specific scoring criteria. It is also a benchmark built by a frontier model vendor that reports its own products’ results on it. That is not an accusation of manipulation. It is a structure, and the structure is worth knowing.

MedHELM is the more interesting case, because it looks like the independent alternative. It originated at Stanford, through the Center for Research on Foundation Models and Stanford Health Care, in collaboration with Microsoft Healthcare and Life Sciences. Its own site states that in 2026 it became an independent, community-led project under an Apache 2.0 licence.

The same page states that Pacific AI provides technical stewardship of the codebase. Pacific AI is a commercial healthcare AI governance company. Its platform sells automated model testing pipelines that explicitly include MedHELM benchmarks. It is also a certified Assurance Resource Provider with the Coalition for Health AI, the industry body that accredits the vendors who audit clinical AI. The community-owned code lives in Pacific AI’s GitHub organization rather than a university or foundation namespace.

So the steward of the benchmark sells the compliance product that runs the benchmark, and is accredited by the coalition that governs the benchmark’s users. Every link in that chain is publicly documented on the parties’ own websites. None of it is hidden. It is simply that nobody had assembled it in one place, and “independent, community-led” is doing a great deal of work in the sentence where it appears.

How current are healthcare AI leaderboard scores?

Older than almost anyone assumes, and the staleness is structurally invisible.

Featured Partner

Invest in the Infrastructure Behind Modern Medicine

As healthcare expands beyond hospital walls, the buildings and campuses supporting that shift are generating compelling returns for investors who move early. The Healthcare Real Estate Fund offers qualified investors direct access to a curated portfolio of medical office, outpatient, and specialty care facilities.

Learn More →

The MedHELM leaderboard is at version 5.0.0, last updated 14 May 2026. As of 5 August 2026 that is 83 days. It ranks 11 models. Among those 11 it places a 2026 frontier model and a February 2025 model in the same ranked column under a single mean win rate, which inflates the apparent spread between them and rewards recency rather than quality.

The deeper problem is that leaderboards publish current state only. When a board updates, the previous version is overwritten and gone. There is no public record of what MedHELM version 4 said, which means nobody can audit how rankings have drifted, whether a model quietly fell, or whether the methodology changed underneath the numbers. We now snapshot these boards so that history exists going forward. It does not exist looking backward.

Can anyone reproduce these results?

Often not. MedHELM draws on 31 datasets. Thirteen are public, six are gated behind an approval process, and twelve are private. That means 41.9 percent of the benchmark can be independently reproduced by a third party, verified 5 August 2026. The project offers seven non-gated datasets as substitutes for the gated ones, which is a reasonable accommodation and also an admission.

MedHELM grades open-ended answers using an ensemble of language models acting as a jury. Its published agreement with expert clinicians is an intraclass correlation of 0.47, and the project notes that this exceeds clinician-to-clinician agreement on the same tasks. Both statements are true. By conventional interpretation an intraclass correlation below 0.50 is poor, and the honest reading is that the underlying ground truth is noisy, not that the automated jury is good.

The ACL 2026 audit that seeded our registry found this pattern across the field. Of 53 medical benchmarks examined, 92 percent had no mechanism to detect or handle training data contamination, 94 percent did not test model robustness, 96 percent did not evaluate whether a model can express uncertainty, and 85 percent had no stated plan for long-term maintenance.

Which leaderboards are estimating rather than measuring?

At least one popular medical model leaderboard states on its own page that several of its frontier model figures are approximate, derived from early evaluation reports and extrapolated from published general reasoning scores. Those estimates appear in a ranked table alongside measured results, formatted identically.

We now record this as a field. Every entity in the registry is marked as measured, self-reported, estimated or extrapolated. It is a distinction that takes one column to capture and changes what a number means entirely.

The case that shows why any of this matters

Benchmarks would be an academic curiosity if they stayed in papers. They do not.

OpenEvidence is a clinical AI reference tool that, according to figures the company gave NBC News, was used by about 65 percent of United States doctors across nearly 27 million clinical encounters in April 2026 alone. It is free to physicians. Its revenue comes from pharmaceutical and medical device advertising shown to clinicians.

On 12 June 2026, Nature Medicine published a brief communication from a team at NYU Langone comparing two specialized clinical AI tools, OpenEvidence and UpToDate Expert AI, against three general frontier models: GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6. The evaluation ran in three stages: 500 medical licensing exam questions, 500 HealthBench items, and a benchmark of 100 real de-identified physician queries reviewed by 12 blinded clinicians producing 1,800 annotations.

The frontier models won all three stages. On the exam questions, Gemini scored 97.4 percent, GPT 94.2 and Claude 90.2, against 89.6 for OpenEvidence and 88.4 for UpToDate Expert AI. On the real physician queries, the paper reported that the specialized clinical tools performed comparably to Google’s automatically enabled search summaries.

The study has a genuine methodological weakness, and it belongs in any honest account. The frontier models were run through clean programming interfaces at a fixed temperature with search enabled, while the two clinical tools were queried by hand through their consumer web interfaces. That is a real confounder, and critics raised it immediately.

What happened next is the part that belongs in an article about who grades the test. Rather than file a letter to the editor or a rebuttal paper, the company responded publicly, alleging benchmark contamination, misrepresented metrics and an undisclosed conflict of interest. Then, according to reporting by STAT, it approached a University of California, San Francisco researcher to conduct a new study, supplied the data, paid the survey respondents and helped develop the data collection plan. The resulting preprint, built on 620 real point-of-care queries and graded by 149 practicing physicians matched to the relevant specialty, found the specialized tool scored highest on every dimension measured. STAT reported that the company’s involvement was not disclosed in the study and that its executives are not listed as authors.

Both studies cannot be the whole truth, and neither is simply wrong. That is the point. A company accused academics of an undisclosed conflict of interest, then funded, staffed and supplied a study that reached the opposite conclusion, without disclosing that it had done so. We are not characterizing intent. We are describing a structure, and the structure answers the question this article opened with. Whoever controls the test controls the result, and in healthcare AI right now the people who control the tests are frequently the people being tested.

What this means for you

If you are choosing a model for clinical work, three things follow.

Start from the deployment, not the ranking. Exam benchmarks tell you almost nothing about clinical documentation. A safety benchmark matters more than any accuracy score before a patient-facing launch. Each entry in our registry states plainly what it is for and what it should not be used for.

Check who owns the scoreboard before you trust the score. Not because ownership makes a benchmark worthless, but because it tells you which failure modes to look for.

Check the date. A leaderboard three months behind will rank a model you cannot buy above one you can.

How we verified this, and what we refuse to publish

Every governance claim in the registry carries a linked primary source and a verification date, and is labelled verified, supported, disputed or unknown. Nothing is inferred from a company name or a press release. Where an instrument has no recorded conflict, we state explicitly that none has been verified rather than implying none exists.

We publish no single independence score. A composite ranking of other people’s benchmarks, built on a rubric we have not validated for inter-rater reliability, would be exactly the failure this article documents. When we have run that validation, we will publish the agreement statistics alongside the scores. Not before.

We also disclose our own position. We operate none of the instruments listed and accept no payment from any entity in the registry. We do develop a separate methodology for appraising the evidence behind healthcare AI claims, which evaluates claims and companies rather than models, and which has not yet published its own reliability data. A registry about undisclosed interests should begin by disclosing its own.

Frequently asked questions

Who regulates healthcare AI benchmarks?

No one. Medical AI benchmarks are published by universities, model vendors, commercial governance companies and independent researchers, with no accreditation requirement and no common standard. Certification bodies such as the Coalition for Health AI accredit the vendors who audit clinical AI, but they do not govern the benchmarks themselves.

Are healthcare AI benchmarks independent?

Several are not fully independent. Some are created by model vendors that report their own results, and at least one major benchmark is maintained by a commercial company that sells a governance product depending on it. Every entity page in our registry records the creator, funder and steward with a linked source.

Can medical AI benchmark scores be compared across leaderboards?

No. A mean win rate, an Elo rating and a percent accuracy are different units. Two leaderboards can rank the same models in opposite orders without either being wrong.

Which healthcare AI benchmark is the best one?

There is no single best benchmark, because they measure different things. Match the instrument to the deployment: exam benchmarks for knowledge comparison, task benchmarks for clinical workflows, safety benchmarks before a patient-facing launch, and agentic benchmarks for tools that take actions inside a record system.

How often are healthcare AI leaderboards updated?

Inconsistently, and staleness is usually invisible because leaderboards publish current state only and discard version history. The most substantial clinical benchmark was 83 days past its last update as of 5 August 2026.

What is MedCheck?

MedCheck is a peer-reviewed lifecycle audit framework published at ACL 2026 that scored 53 medical benchmarks against 46 criteria across five phases, from design through governance. It measures benchmark quality but does not record ownership or funding, which is the layer our registry adds.

The full registry, with all 63 entities, their ownership chains and their verification dates, is at the Healthcare AI Benchmark Registry. If you want the deeper explanation of how one of these scoreboards is actually built, start with our breakdown of what HealthBench measures and what it cannot.

Free Daily Briefing

The Latest Longevity Science.
Delivered Every Morning.

Join researchers, physicians, and health professionals getting daily breakthroughs in AI-driven medicine, epigenetics, and longevity research.

Support the research that powers this editorial

No spam. Unsubscribe anytime. We respect your inbox.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *