Observed arrival · 2026-10-09
Benchgap estimates the missing scores in LLM leaderboards
A public LLM benchmark leaderboard that distinguishes measured results from estimated ones and attaches error ranges and confidence levels to estimates.
- For
- LLM evaluators comparing model benchmark results
- Worth noticing
- Estimated entries include cross-validated error ranges and confidence labels, rather than being presented as measured results.
Field notes
Each leaderboard separates measured results from estimates, and estimated rows carry an error range and confidence label. The homepage reports coverage of 402 models across 148 benchmark versions, with 4,919 measured scores and 16,584 estimates; it also identifies public leaderboard sources including Artificial Analysis, Published, and Vals AI. Its navigation exposes calibration, multivariate analysis, harness tax, method, publications, and an API, though the supplied extract does not show how those sections work.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue