A very large multi-agent generated reasoning dataset for medical training and evaluation.
Audited by MedCheckSun et al., 2025
Key facts
| What it measures | A very large multi-agent generated reasoning dataset for medical training and evaluation. |
| Who should care | Teams fine-tuning for medical reasoning. |
| Use it when | Training data and reasoning evaluation at scale. |
| Do not use it for | Independent evaluation. It is largely model-generated, so contamination and circularity risk are real. |
| Score scale | Reasoning accuracy on the generated set. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | research |
| Citation | Sun et al., 2025 |
| Status | active |
Frequently asked questions
What does ReasonMed measure?
A very large multi-agent generated reasoning dataset for medical training and evaluation.
Who should use ReasonMed?
Teams fine-tuning for medical reasoning. Training data and reasoning evaluation at scale.
What should ReasonMed not be used for?
Independent evaluation. It is largely model-generated, so contamination and circularity risk are real.
How are ReasonMed scores reported?
Reasoning accuracy on the generated set. Scores from different benchmarks are not comparable to each other.
Is ReasonMed independent?
No conflicts of interest have been verified for ReasonMed in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Chinese medical question and answer selection built from online health forums. One of the oldest entries in th
Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished
The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Long-context medical evaluation running up to 200,000 tokens.
