Long-context medical evaluation running up to 200,000 tokens.
Audited by MedCheckFan et al., NAACL Findings 2025
Key facts
| What it measures | Long-context medical evaluation running up to 200,000 tokens. |
| Who should care | Teams feeding whole charts or long documents into a model. |
| Use it when | Testing whether performance survives a long context. |
| Do not use it for | Short prompt evaluation. |
| Score scale | Accuracy measured at increasing context lengths. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English, Chinese |
| Use cases | documentation, research |
| Citation | Fan et al., NAACL Findings 2025 |
| Status | active |
Frequently asked questions
What does MedOdyssey measure?
Long-context medical evaluation running up to 200,000 tokens.
Who should use MedOdyssey?
Teams feeding whole charts or long documents into a model. Testing whether performance survives a long context.
What should MedOdyssey not be used for?
Short prompt evaluation.
How are MedOdyssey scores reported?
Accuracy measured at increasing context lengths. Scores from different benchmarks are not comparable to each other.
Is MedOdyssey independent?
No conflicts of interest have been verified for MedOdyssey in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse
Whether a model can find and fix clinical errors already present in a note.
