Long-context medical evaluation running up to 200,000 tokens.

MedOdyssey is a benchmark in healthcare AI. Long-context medical evaluation running up to 200,000 tokens. Scores are reported as: Accuracy measured at increasing context lengths.

Audited by MedCheckFan et al., NAACL Findings 2025

Key facts

What it measuresLong-context medical evaluation running up to 200,000 tokens.
Who should careTeams feeding whole charts or long documents into a model.
Use it whenTesting whether performance survives a long context.
Do not use it forShort prompt evaluation.
Score scaleAccuracy measured at increasing context lengths.
Grader methodunknown
First released2025
LanguagesEnglish, Chinese
Use casesdocumentation, research
CitationFan et al., NAACL Findings 2025
Statusactive

Frequently asked questions

What does MedOdyssey measure?

Long-context medical evaluation running up to 200,000 tokens.

Who should use MedOdyssey?

Teams feeding whole charts or long documents into a model. Testing whether performance survives a long context.

What should MedOdyssey not be used for?

Short prompt evaluation.

How are MedOdyssey scores reported?

Accuracy measured at increasing context lengths. Scores from different benchmarks are not comparable to each other.

Is MedOdyssey independent?

No conflicts of interest have been verified for MedOdyssey in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedConceptsQA

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse

MEDEC

Whether a model can find and fix clinical errors already present in a note.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities