Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data.
Audited by MedCheckOlatunji et al., 2025
Key facts
| What it measures | Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data. |
| Who should care | Anyone deploying in low and middle income settings, or worried a model only knows North American medicine. |
| Use it when | Testing geographic and demographic generalization. |
| Do not use it for | Comparing against the standard Western exam literature. The question set is deliberately different. |
| Score scale | Percent accuracy, multiple choice and short answer. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English, French, other |
| Use cases | exam knowledge, equity |
| Citation | Olatunji et al., 2025 |
| Status | active |
Frequently asked questions
What does AfriMed-QA measure?
Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data.
Who should use AfriMed-QA?
Anyone deploying in low and middle income settings, or worried a model only knows North American medicine. Testing geographic and demographic generalization.
What should AfriMed-QA not be used for?
Comparing against the standard Western exam literature. The question set is deliberately different.
How are AfriMed-QA scores reported?
Percent accuracy, multiple choice and short answer. Scores from different benchmarks are not comparable to each other.
Is AfriMed-QA independent?
No conflicts of interest have been verified for AfriMed-QA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Chinese-language evaluation of physical and mental health knowledge in large models.
Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.
Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.
Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic
Spanish healthcare specialization exam questions requiring multi-step reasoning.
Indian medical entrance exam questions across many subjects. Very large and very widely used.
