A very large multi-agent generated reasoning dataset for medical training and evaluation.

ReasonMed is a benchmark in healthcare AI. A very large multi-agent generated reasoning dataset for medical training and evaluation. Scores are reported as: Reasoning accuracy on the generated set.

Audited by MedCheckSun et al., 2025

Key facts

What it measuresA very large multi-agent generated reasoning dataset for medical training and evaluation.
Who should careTeams fine-tuning for medical reasoning.
Use it whenTraining data and reasoning evaluation at scale.
Do not use it forIndependent evaluation. It is largely model-generated, so contamination and circularity risk are real.
Score scaleReasoning accuracy on the generated set.
Grader methodunknown
First released2025
LanguagesEnglish
Use casesresearch
CitationSun et al., 2025
Statusactive

Paper

Frequently asked questions

What does ReasonMed measure?

A very large multi-agent generated reasoning dataset for medical training and evaluation.

Who should use ReasonMed?

Teams fine-tuning for medical reasoning. Training data and reasoning evaluation at scale.

What should ReasonMed not be used for?

Independent evaluation. It is largely model-generated, so contamination and circularity risk are real.

How are ReasonMed scores reported?

Reasoning accuracy on the generated set. Scores from different benchmarks are not comparable to each other.

Is ReasonMed independent?

No conflicts of interest have been verified for ReasonMed in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CLIMB

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

cMedQA2

Chinese medical question and answer selection built from online health forums. One of the oldest entries in th

DataDEL

Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished

HELM

The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedOdyssey

Long-context medical evaluation running up to 200,000 tokens.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities