Whether a model computes clinical scores and formulas correctly from a patient note.

MedCalc-Bench is a benchmark in healthcare AI. Whether a model computes clinical scores and formulas correctly from a patient note. Scores are reported as: Percent accuracy on computed values.

Audited by MedCheckKhandekar et al., NeurIPS 2024

Key facts

What it measuresWhether a model computes clinical scores and formulas correctly from a patient note.
Who should careAnyone deploying a model anywhere near dosing or risk scoring.
Use it whenYou need arithmetic and formula reliability, which is where models quietly fail.
Do not use it forConversational or reasoning quality.
Score scalePercent accuracy on computed values.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesclinical decision support, safety
CitationKhandekar et al., NeurIPS 2024
Statusactive

Paper

Frequently asked questions

What does MedCalc-Bench measure?

Whether a model computes clinical scores and formulas correctly from a patient note.

Who should use MedCalc-Bench?

Anyone deploying a model anywhere near dosing or risk scoring. You need arithmetic and formula reliability, which is where models quietly fail.

What should MedCalc-Bench not be used for?

Conversational or reasoning quality.

How are MedCalc-Bench scores reported?

Percent accuracy on computed values. Scores from different benchmarks are not comparable to each other.

Is MedCalc-Bench independent?

No conflicts of interest have been verified for MedCalc-Bench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities