Observed arrival · 2026-09-05
AIBench, an AI leaderboard that explains its homework
A benchmark index that ranks AI models across categories, compares capability with token price, and exposes the methodology behind each number.
Field notes
The ranking uses a two-parameter item-response model rather than hand-picked benchmark weights: each benchmark receives estimated difficulty and discrimination, and saturated or noisy tests contribute less. Category scores are then shrunk toward a pooled estimate when evidence is thin, while equal category weighting prevents one skill from dominating. The page also exposes a 90% bootstrap interval and a fragility statistic based on dropping one active benchmark at a time.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue