Bilingual medical visual question answering with a semantic knowledge layer attached to the images.
Audited by MedCheckLiu et al., IEEE ISBI 2021
Key facts
| What it measures | Bilingual medical visual question answering with a semantic knowledge layer attached to the images. |
| Who should care | Vision-language teams working across English and Chinese. |
| Use it when | Knowledge-grounded medical VQA. |
| Do not use it for | Text-only or English-only evaluation. |
| Score scale | Percent accuracy, open and closed questions. |
| Grader method | unknown |
| First released | 2021 |
| Languages | English, Chinese |
| Use cases | imaging, multimodal |
| Citation | Liu et al., IEEE ISBI 2021 |
| Status | active |
Frequently asked questions
What does SLAKE measure?
Bilingual medical visual question answering with a semantic knowledge layer attached to the images.
Who should use SLAKE?
Vision-language teams working across English and Chinese. Knowledge-grounded medical VQA.
What should SLAKE not be used for?
Text-only or English-only evaluation.
How are SLAKE scores reported?
Percent accuracy, open and closed questions. Scores from different benchmarks are not comparable to each other.
Is SLAKE independent?
No conflicts of interest have been verified for SLAKE in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Multimodal evaluation for endoscopy image and video analysis.
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
