Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.

MedConceptsQA is a benchmark in healthcare AI. Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes. Scores are reported as: Percent accuracy across code systems and difficulty tiers.

Audited by MedCheckShoham and Rappoport, Computers in Biology and Medicine 2024

Key facts

What it measuresWhether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.
Who should careRevenue cycle, coding and claims teams.
Use it whenEvaluating coding and terminology reliability.
Do not use it forClinical reasoning or patient communication.
Score scalePercent accuracy across code systems and difficulty tiers.
Grader methodunknown
First released2024
LanguagesEnglish
Use casesadministration, documentation
CitationShoham and Rappoport, Computers in Biology and Medicine 2024
Statusactive

Frequently asked questions

What does MedConceptsQA measure?

Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.

Who should use MedConceptsQA?

Revenue cycle, coding and claims teams. Evaluating coding and terminology reliability.

What should MedConceptsQA not be used for?

Clinical reasoning or patient communication.

How are MedConceptsQA scores reported?

Percent accuracy across code systems and difficulty tiers. Scores from different benchmarks are not comparable to each other.

Is MedConceptsQA independent?

No conflicts of interest have been verified for MedConceptsQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

MedHELM

Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst

BRIDGE

Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task

ClinicBench

Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.

EHRNoteQA

Questions answered from real discharge summaries, written by clinicians for real-world practice.

MedAgentBench

A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n

MEDEC

Whether a model can find and fix clinical errors already present in a note.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities