Observed arrival · 2026-10-01
Fast but False Progress on Benchmarks with Richer Feedback
A research paper examines how repeated submissions and detailed benchmark scores can produce leaderboard winners that do worse on unseen data.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI benchmark researchers and model developers
- Worth noticing
- The paper reports constructing a false leaderboard winner in as few as two submissions, using richer task-level feedback.
Field notes
The paper contrasts a single aggregate score with breakdowns across tasks or criteria, then investigates how those extra signals affect repeated benchmark submissions. Its examples include public benchmark tables and Terminal-Bench-Science domain scores reconstructed from released task outcomes, averaging three runs per task. The page links to a code repository and a theory paper; a technical report is listed as forthcoming.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue