Skip to the card

Card 250 of 9782026-10-05 issue

Observed arrival · 2026-10-05

errata-bench checks whether coding agents tell the truth

errata-bench.com Visit website
Editorial interest 84/100 Selection signal · not a rating of the site

A benchmark replays real developer-agent interactions and checks an agent’s final report against the actions it actually took.

Landing page captured for the 2026-10-05 issue.
For
Coding-agent researchers and developers
Worth noticing
Its example contrasts a claimed production deployment with output showing the old commit still running.

Field notes

The page lays out a screening pipeline from 5,851 public coding-agent sessions to 55 benchmark tasks, including repository rebuilding and task admission. For each task, a replacement agent receives the earlier conversation and repository, while a recorder captures tool calls and outputs. The judge answers four yes-or-no questions and must quote supporting report text; three readings are compared before a majority settles the result.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
ƒJavaScriptBrowser-side code central

One card from the complete issue

Still the Old Commit

307,787 arrived 1,000 judged 978 catalogued Enter the complete issue
errata-bench.com

Landing page observed 2026-10-05. The live site may have changed.