Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.
Audited by MedCheckWu et al., 2025
Key facts
| What it measures | Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks. |
| Who should care | Health systems that care whether a model handles their own messy documentation. |
| Use it when | You want the closest available proxy for real clinical text. |
| Do not use it for | Quick comparison. It is large and heterogeneous. |
| Score scale | Task-specific metrics, reported per task rather than as one number. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English, Chinese, other |
| Use cases | documentation, clinical decision support |
| Citation | Wu et al., 2025 |
| Status | active |
Frequently asked questions
What does BRIDGE measure?
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and tasks.
Who should use BRIDGE?
Health systems that care whether a model handles their own messy documentation. You want the closest available proxy for real clinical text.
What should BRIDGE not be used for?
Quick comparison. It is large and heterogeneous.
How are BRIDGE scores reported?
Task-specific metrics, reported per task rather than as one number. Scores from different benchmarks are not comparable to each other.
Is BRIDGE independent?
No conflicts of interest have been verified for BRIDGE in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confuse
Whether a model can find and fix clinical errors already present in a note.
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
