Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated moment.

MedJourney is a benchmark in healthcare AI. Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated moment. Scores are reported as: Stage-level metrics across the journey.

Audited by MedCheckWu et al., NeurIPS 2024

Key facts

What it measuresFollows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated moment.
Who should careCare coordination and longitudinal care product teams.
Use it whenEvaluating performance across care stages.
Do not use it forSingle-encounter tasks.
Score scaleStage-level metrics across the journey.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesclinical decision support, patient communication
CitationWu et al., NeurIPS 2024
Statusactive

Frequently asked questions

What does MedJourney measure?

Follows a patient across the full clinical journey, from first contact through follow-up, rather than one isolated moment.

Who should use MedJourney?

Care coordination and longitudinal care product teams. Evaluating performance across care stages.

What should MedJourney not be used for?

Single-encounter tasks.

How are MedJourney scores reported?

Stage-level metrics across the journey. Scores from different benchmarks are not comparable to each other.

Is MedJourney independent?

No conflicts of interest have been verified for MedJourney in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities