Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions.

DataDEL is a benchmark in healthcare AI. Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions. Scores are reported as: Data efficiency metrics, not answer accuracy.

Audited by MedCheckYang et al., ACM Multimedia 2024

Key facts

What it measuresMeasures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions.
Who should careTeams pretraining medical foundation models.
Use it whenAssessing data efficiency and curation strategy.
Do not use it forChoosing between off-the-shelf models. It measures a different thing entirely.
Score scaleData efficiency metrics, not answer accuracy.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesresearch
CitationYang et al., ACM Multimedia 2024
Statusactive

Frequently asked questions

What does DataDEL measure?

Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions.

Who should use DataDEL?

Teams pretraining medical foundation models. Assessing data efficiency and curation strategy.

What should DataDEL not be used for?

Choosing between off-the-shelf models. It measures a different thing entirely.

How are DataDEL scores reported?

Data efficiency metrics, not answer accuracy. Scores from different benchmarks are not comparable to each other.

Is DataDEL independent?

No conflicts of interest have been verified for DataDEL in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CLIMB

Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.

cMedQA2

Chinese medical question and answer selection built from online health forums. One of the oldest entries in th

HELM

The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

MedOdyssey

Long-context medical evaluation running up to 200,000 tokens.

MedS-Bench

Evaluates models across eleven clinical task types beyond question answering, including summarization, extract

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities