Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical predecessor that MedCheck adapted.
Reuel et al., NeurIPS 2024
Key facts
| What it measures | Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical predecessor that MedCheck adapted. |
| Who should care | Anyone judging whether a benchmark is well built before trusting its scores. |
| Use it when | Evaluating benchmark quality outside medicine. |
| Do not use it for | Anything medical, and any model ranking. It has none. |
| Score scale | Percent of checklist criteria met. |
| Grader method | human expert |
| First released | 2024 |
| Languages | English |
| Use cases | meta |
| Citation | Reuel et al., NeurIPS 2024 |
| Status | active |
Independence assessment
Stanford. General purpose, not medical specific.
Scope
General-purpose AI benchmark assessment framework, 46 criteria. The non-medical predecessor MedCheck adapted.
Frequently asked questions
What does BetterBench measure?
Assesses general-purpose AI benchmarks themselves against a 46-item best-practice checklist. The non-medical predecessor that MedCheck adapted.
Who should use BetterBench?
Anyone judging whether a benchmark is well built before trusting its scores. Evaluating benchmark quality outside medicine.
What should BetterBench not be used for?
Anything medical, and any model ranking. It has none.
How are BetterBench scores reported?
Percent of checklist criteria met. Scores from different benchmarks are not comparable to each other.
Is BetterBench independent?
No conflicts of interest have been verified for BetterBench in this registry as of 2026-08-05. Absence of a recorded conflict means none has been verified, not that none exists.
Compare with
Audits the benchmarks themselves against a 46-item lifecycle checklist. Does not rank models.
