What we record, where it comes from, and what we refuse to publish.

The registry scores evaluation instruments against 56 criteria: the 46 published by MedCheck, adopted verbatim, plus 10 covering ownership, incentives and independence that no existing framework measures. Every claim carries a linked source and a verification date.

The scoring instrument

Phase 1. Design and Conceptualization MedCheck, adopted verbatim

  1. Does the benchmark define the targeted LLM capabilities in medicine for evaluation (e.g., QA, diagnostic reasoning)?
  2. Does it describe specific clinical or research applications and their potential value?
  3. Does it highlight its unique contribution or innovation compared to existing benchmarks?
  4. Does it define the specific LLM functions for evaluation?
  5. Does it define the scope of the medical specialties of the benchmark?
  6. Is it designed to meet the needs of LLM researchers and clinical challenges?
  7. Have qualified medical experts been involved in benchmark development?
  8. Does it rely on recognized medical sources (e.g., clinical guidelines, databases)?
  9. Does it align with international medical standards (e.g., ICD, SNOMED CT, LOINC)?
  10. Are the metrics clear and closely related to clinical tasks?
  11. Does it assess multiple dimensions (e.g., safety, completeness, interpretability)?
  12. Does it consider potential risks and biases in model outputs?

Phase 2. Dataset Construction and Management MedCheck, adopted verbatim

  1. Are the original sources clearly stated and traceable?
  2. Are the sources authoritative and well-justified?
  3. Is the data origin clear, and is synthetic data validated?
  4. Is the dataset representative of the target population?
  5. Are diversity goals defined with supporting quantitative analysis?
  6. Is the data properly cleaned and standardized?
  7. Are privacy measures described and regulation-compliant?
  8. Is the data format consistent and unambiguous?
  9. Is expert-involved data review in place?
  10. Are reference answers accurate and validated?
  11. Are contamination risks detected and handled?

Phase 3. Technical Implementation and Evaluation Methodology MedCheck, adopted verbatim

  1. Is the evaluation tool easy to install and use?
  2. Are detailed documentation and environment settings provided to support reproducibility?
  3. Are baseline models or human performance results provided for comparison?
  4. Are there evaluations for the model reasoning process?
  5. Are there evaluations testing the model robustness (e.g., input perturbations)?
  6. Does the benchmark design help evaluate the generalization ability of models to unseen data?
  7. Are there evaluations testing the model ability to express uncertainty?
  8. Does the benchmark support both closed-source APIs and open-source models?

Phase 4. Benchmark Validity and Performance Verification MedCheck, adopted verbatim

  1. Does the benchmark cover the claimed medical knowledge and skills?
  2. Do the tasks realistically simulate clinical settings?
  3. Can it distinguish between models of different levels?
  4. Are benchmark scores correlated with real-world clinical performance?
  5. Does the benchmark demonstrate internal consistency to ensure that different components reliably assess the same capability?
  6. Are statistical tests used to verify results?

Phase 5. Documentation, Openness and Governance MedCheck, adopted verbatim

  1. Is there clear, comprehensive benchmark documentation?
  2. Are the evaluation criteria and instructions clear and easy to follow?
  3. Are limitations and potential risks openly discussed?
  4. Has the benchmark undergone formal academic peer review?
  5. Are the code and data publicly available with proper licensing?
  6. Is a clear usage and citation guideline provided?
  7. Is there a clear plan for updates and version control?
  8. Is there a public channel for user feedback?
  9. Is the long-term maintenance responsibility clearly stated?

Phase 6. Ownership, Incentives and Independence Healthcare Discovery addition

  1. Is the funding source of the benchmark disclosed by name?
  2. Is the current code steward identified, and is the steward independent of model vendors?
  3. Does the steward, host or maintainer sell a commercial product that uses or depends on this benchmark?
  4. Do any creators, funders or stewards also have models ranked on its own leaderboard?
  5. Is the leaderboard updated on a stated cadence, and is the last update within that cadence?
  6. Does the leaderboard avoid ranking models from different release generations under a single composite metric?
  7. Can a third party fully reproduce the published leaderboard from public artifacts alone?
  8. Are published scores measured by the operator rather than estimated, extrapolated or self-reported?
  9. Is there a stated conflict-of-interest policy governing model submissions and vendor participation?
  10. Is the governing body distinct from any body that certifies, accredits or sells services to the benchmark users?

Evidence rules

No claim is recorded without a retrievable source and a date. Every claim is labelled verified, supported, disputed or unknown. Absence of a recorded conflict means none has been verified, not that none exists, and every entity page says so explicitly.

The editorial posture is referee, not plaintiff. Where a conflict exists we describe the structure and let the reader judge. We do not characterize intent.

What we do not publish

No composite independence score until the ownership criteria have been scored across the registry and the agreement between scorers has been measured. No estimated or extrapolated numbers presented as measurements. No leaderboard rank comparison across instruments that use different scales.

Our own incentives

Healthcare Discovery is an independent editorial platform. We do not operate any of the instruments listed here, and we accept no payment from any entity in this registry.

We do develop a separate methodology, the HYPE Index, for appraising the evidence behind healthcare AI claims. It evaluates claims, papers and companies rather than models, so it is not a benchmark, it is not listed in this registry, and it competes with nothing here. We disclose it anyway, because a registry about undisclosed interests should begin by disclosing its own. The HYPE Index has not yet published inter-rater reliability data, which is the same standard we apply to every instrument on this site.

Corrections from the operators of any listed instrument are welcome and will be published with the correction date.

Frequently asked questions

How is this registry verified?

Every governance claim points at a retrieved primary source with a date. Claims are labelled verified, supported, disputed or unknown. Nothing is inferred from a company name or a press release alone.

Why is there no overall independence score?

Because the scoring has not been run and validated yet. Publishing an unvalidated composite that ranks other people’s benchmarks would repeat the exact failure this registry documents.

Where does the list of benchmarks come from?

The seed inventory is the set audited by MedCheck (Ma et al., ACL 2026). Living leaderboards, arenas, aggregators and certification bodies were added because MedCheck’s scope excluded them.

Last updated 2026-08-05. Back to the registry