Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic questions.
Audited by MedCheckPanagoulias et al., 2024
Key facts
| What it measures | Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic questions. |
| Who should care | Researchers looking for a specialty-segmented question set. |
| Use it when | Specialty-level knowledge probes. |
| Do not use it for | Workflow or documentation evaluation. |
| Score scale | Percent accuracy by specialty. |
| Grader method | unknown |
| First released | 2024 |
| Languages | English |
| Use cases | exam knowledge |
| Citation | Panagoulias et al., 2024 |
| Status | active |
Frequently asked questions
What does COGNET-MD measure?
Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic questions.
Who should use COGNET-MD?
Researchers looking for a specialty-segmented question set. Specialty-level knowledge probes.
What should COGNET-MD not be used for?
Workflow or documentation evaluation.
How are COGNET-MD scores reported?
Percent accuracy by specialty. Scores from different benchmarks are not comparable to each other.
Is COGNET-MD independent?
No conflicts of interest have been verified for COGNET-MD in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Multi-specialty medical questions written by clinicians and students across African countries, built to test w
Chinese-language evaluation of physical and mental health knowledge in large models.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.
Spanish healthcare specialization exam questions requiring multi-step reasoning.
Indian medical entrance exam questions across many subjects. Very large and very widely used.
