Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.
Audited by MedCheckShoham and Rappoport, Computers in Biology and Medicine 2024
Key facts
| What it measures | Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes. |
| Who should care | Revenue cycle, coding and claims teams. |
| Use it when | Evaluating coding and terminology reliability. |
| Do not use it for | Clinical reasoning or patient communication. |
| Score scale | Percent accuracy across code systems and difficulty tiers. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | administration, documentation |
| Citation | Shoham and Rappoport, Computers in Biology and Medicine 2024 |
| Status | active |
Frequently asked questions
What does MedConceptsQA measure?
Whether a model understands medical coding vocabularies such as ICD and ATC, including rare and easily confused codes.
Who should use MedConceptsQA?
Revenue cycle, coding and claims teams. Evaluating coding and terminology reliability.
What should MedConceptsQA not be used for?
Clinical reasoning or patient communication.
How are MedConceptsQA scores reported?
Percent accuracy across code systems and difficulty tiers. Scores from different benchmarks are not comparable to each other.
Is MedConceptsQA independent?
No conflicts of interest have been verified for MedConceptsQA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Tests models on 121 real clinical work tasks, not exam questions: summarizing records, drafting discharge inst
Real clinical practice text, drawn from actual notes rather than exam material, across many languages and task
Broad clinical benchmark covering summarization, diagnosis and treatment planning in one suite.
Questions answered from real discharge summaries, written by clinicians for real-world practice.
A realistic virtual electronic health record environment where agents must complete tasks by taking actions, n
Whether a model can find and fix clinical errors already present in a note.
