Real-world clinical benchmark with physician validation of both questions and reference answers.

LLMEval-Med is a benchmark in healthcare AI. Real-world clinical benchmark with physician validation of both questions and reference answers. Scores are reported as: Physician-validated scoring on open responses.

Audited by MedCheckZhang et al., 2025

Key facts

What it measuresReal-world clinical benchmark with physician validation of both questions and reference answers.
Who should careTeams that want clinician-validated ground truth rather than scraped exam keys.
Use it whenYou care that a doctor checked the answer key.
Do not use it forLarge-scale automated comparison. Physician validation limits the size.
Score scalePhysician-validated scoring on open responses.
Grader methodunknown
First released2025
LanguagesChinese, English
Use casesclinical decision support
CitationZhang et al., 2025
Statusactive

Paper

Frequently asked questions

What does LLMEval-Med measure?

Real-world clinical benchmark with physician validation of both questions and reference answers.

Who should use LLMEval-Med?

Teams that want clinician-validated ground truth rather than scraped exam keys. You care that a doctor checked the answer key.

What should LLMEval-Med not be used for?

Large-scale automated comparison. Physician validation limits the size.

How are LLMEval-Med scores reported?

Physician-validated scoring on open responses. Scores from different benchmarks are not comparable to each other.

Is LLMEval-Med independent?

No conflicts of interest have been verified for LLMEval-Med in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

AgentClinic

Simulated clinical encounters where the model must gather information over several turns rather than answer on

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

CliMedBench

Large-scale Chinese benchmark built around real clinical scenarios rather than exam questions.

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities