Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.

BRIDGE is a benchmark in healthcare AI. Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks. Scores are reported as: Task-specific metrics, reported per task rather than as one number.

Audited by MedCheckWu et al., 2025

Key facts

What it measuresReal clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.
Who should careHealth systems that care whether a model handles their own messy documentation.
Use it whenYou want the closest available proxy for real clinical text.
Do not use it forQuick comparison. It is large and heterogeneous.
Score scaleTask-specific metrics, reported per task rather than as one number.
Grader methodunknown
First released2025
LanguagesEnglish, Chinese, other
Use casesdocumentation, clinical decision support
CitationWu et al., 2025
Statusactive

Paper

Frequently asked questions

What does BRIDGE measure?

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.

Who should use BRIDGE?

Health systems that care whether a model handles their own messy documentation. You want the closest available proxy for real clinical text.

What should BRIDGE not be used for?

Quick comparison. It is large and heterogeneous.

How are BRIDGE scores reported?

Task-specific metrics, reported per task rather than as one number. Scores from different benchmarks are not comparable to each other.

Is BRIDGE independent?

No conflicts of interest have been verified for BRIDGE in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MedConceptsQA

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse

MEDEC

Whether a model can find and fix clinical errors already present in a note.

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities