Skip to the card

Card 223 of 9742026-10-01 issue

Observed arrival · 2026-10-01

Fast but False Progress on Benchmarks with Richer Feedback

false-benchmark-progress.com Visit website
Editorial interest 84/100 Selection signal · not a rating of the site

A research paper examines how repeated submissions and detailed benchmark scores can produce leaderboard winners that do worse on unseen data.

Landing page captured for the 2026-10-01 issue.
For
AI benchmark researchers and model developers
Worth noticing
The paper reports constructing a false leaderboard winner in as few as two submissions, using richer task-level feedback.

Field notes

The paper contrasts a single aggregate score with breakdowns across tasks or criteria, then investigates how those extra signals affect repeated benchmark submissions. Its examples include public benchmark tables and Terminal-Bench-Science domain scores reconstructed from released task outcomes, averaging three runs per task. The page links to a code repository and a theory paper; a technical report is listed as forthcoming.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Pink Dolphin Slippers

676,985 arrived 1,000 judged 974 catalogued Enter the complete issue
false-benchmark-progress.com

Landing page observed 2026-10-01. The live site may have changed.