Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data.

AfriMed-QA is a benchmark in healthcare AI. Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data. Scores are reported as: Percent accuracy, multiple choice and short answer.

Audited by MedCheckOlatunji et al., 2025

Key facts

What it measuresMulti-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data.
Who should careAnyone deploying in low and middle income settings, or worried a model only knows North American medicine.
Use it whenTesting geographic and demographic generalization.
Do not use it forComparing against the standard Western exam literature. The question set is deliberately different.
Score scalePercent accuracy, multiple choice and short answer.
Grader methodunknown
First released2025
LanguagesEnglish, French, other
Use casesexam knowledge, equity
CitationOlatunji et al., 2025
Statusactive

Frequently asked questions

What does AfriMed-QA measure?

Multi-specialty medical questions written by clinicians and students across African countries, built to test whether models work outside Western training data.

Who should use AfriMed-QA?

Anyone deploying in low and middle income settings, or worried a model only knows North American medicine. Testing geographic and demographic generalization.

What should AfriMed-QA not be used for?

Comparing against the standard Western exam literature. The question set is deliberately different.

How are AfriMed-QA scores reported?

Percent accuracy, multiple choice and short answer. Scores from different benchmarks are not comparable to each other.

Is AfriMed-QA independent?

No conflicts of interest have been verified for AfriMed-QA in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

CHBench

Chinese-language evaluation of physical and mental health knowledge in large models.

CMB

Comprehensive Chinese medical benchmark spanning exam knowledge and clinical case work.

CMExam

Chinese medical licensing exam questions with expert annotations for reasoning and difficulty.

COGNET-MD

Evaluation framework and dataset covering several medical specialties, built around clinician-style diagnostic

HeadQA

Spanish healthcare specialization exam questions requiring multi-step reasoning.

MedMCQA

Indian medical entrance exam questions across many subjects. Very large and very widely used.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities