Bilingual medical visual question answering with a semantic knowledge layer attached to the images.

SLAKE is a benchmark in healthcare AI. Bilingual medical visual question answering with a semantic knowledge layer attached to the images. Scores are reported as: Percent accuracy, open and closed questions.

Audited by MedCheckLiu et al., IEEE ISBI 2021

Key facts

What it measuresBilingual medical visual question answering with a semantic knowledge layer attached to the images.
Who should careVision-language teams working across English and Chinese.
Use it whenKnowledge-grounded medical VQA.
Do not use it forText-only or English-only evaluation.
Score scalePercent accuracy, open and closed questions.
Grader methodunknown
First released2021
LanguagesEnglish, Chinese
Use casesimaging, multimodal
CitationLiu et al., IEEE ISBI 2021
Statusactive

Frequently asked questions

What does SLAKE measure?

Bilingual medical visual question answering with a semantic knowledge layer attached to the images.

Who should use SLAKE?

Vision-language teams working across English and Chinese. Knowledge-grounded medical VQA.

What should SLAKE not be used for?

Text-only or English-only evaluation.

How are SLAKE scores reported?

Percent accuracy, open and closed questions. Scores from different benchmarks are not comparable to each other.

Is SLAKE independent?

No conflicts of interest have been verified for SLAKE in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.

Compare with

Asclepius

Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.

CheXpert

A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i

EndoBench

Multimodal evaluation for endoscopy image and video analysis.

GMAI-MMBench

Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type

MMMU (Health and Medicine)

The health and medicine subset of a broad expert-level multimodal reasoning benchmark.

OmniMedVQA

Very large medical visual question answering set spanning many modalities and anatomical regions.

Last verified 2026-08-05. Governance and conflict data compiled by Healthcare Discovery from primary sources, each linked above. Seed inventory from Ma et al., Beyond the Leaderboard, ACL 2026. How we verify · All 63 entities