Multimodal evaluation for endoscopy image and video analysis.
Audited by MedCheckLiu et al., 2025
Key facts
| What it measures | Multimodal evaluation for endoscopy image and video analysis. |
| Who should care | Gastroenterology and surgical AI teams. |
| Use it when | Endoscopy-specific model selection. |
| Do not use it for | General clinical evaluation. |
| Score scale | Percent accuracy across endoscopy tasks. |
| Grader method | unknown |
| First released | 2025 |
| Languages | English |
| Use cases | imaging, multimodal |
| Citation | Liu et al., 2025 |
| Status | active |
Frequently asked questions
What does EndoBench measure?
Multimodal evaluation for endoscopy image and video analysis.
Who should use EndoBench?
Gastroenterology and surgical AI teams. Endoscopy-specific model selection.
What should EndoBench not be used for?
General clinical evaluation.
How are EndoBench scores reported?
Percent accuracy across endoscopy tasks. Scores from different benchmarks are not comparable to each other.
Is EndoBench independent?
No conflicts of interest have been verified for EndoBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Spectrum evaluation for multimodal medical models across many imaging modalities and difficulty levels.
A large chest radiograph dataset with uncertainty labels and expert comparison. Predates language models and i
Large multimodal benchmark for general medical AI, spanning many imaging modalities, departments and task type
The health and medicine subset of a broad expert-level multimodal reasoning benchmark.
Very large medical visual question answering set spanning many modalities and anatomical regions.
Expert-level pathology understanding and reasoning from slide images.
