Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions.
Audited by MedCheckYang et al., ACM Multimedia 2024
Key facts
| What it measures | Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions. |
| Who should care | Teams pretraining medical foundation models. |
| Use it when | Assessing data efficiency and curation strategy. |
| Do not use it for | Choosing between off-the-shelf models. It measures a different thing entirely. |
| Score scale | Data efficiency metrics, not answer accuracy. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | research |
| Citation | Yang et al., ACM Multimedia 2024 |
| Status | active |
Frequently asked questions
What does DataDEL measure?
Measures how efficiently a model learns from medical data during pretraining, rather than how well a finished model answers questions.
Who should use DataDEL?
Teams pretraining medical foundation models. Assessing data efficiency and curation strategy.
What should DataDEL not be used for?
Choosing between off-the-shelf models. It measures a different thing entirely.
How are DataDEL scores reported?
Data efficiency metrics, not answer accuracy. Scores from different benchmarks are not comparable to each other.
Is DataDEL independent?
No conflicts of interest have been verified for DataDEL in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Data foundations for large-scale multimodal clinical models, spanning many data types rather than one.
Chinese medical question and answer selection built from online health forums. One of the oldest entries in th
The parent evaluation framework that MedHELM extends. Runs many models across many scenarios on one standardiz
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Long-context medical evaluation running up to 200,000 tokens.
Evaluates models across eleven clinical task types beyond question answering, including summarization, extract
