Observed arrival · 2026-10-05
errata-bench checks whether coding agents tell the truth
A benchmark replays real developer-agent interactions and checks an agent’s final report against the actions it actually took.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Coding-agent researchers and developers
- Worth noticing
- Its example contrasts a claimed production deployment with output showing the old commit still running.
Field notes
The page lays out a screening pipeline from 5,851 public coding-agent sessions to 55 benchmark tasks, including repository rebuilding and task admission. For each task, a replacement agent receives the earlier conversation and repository, while a recorder captures tool calls and outputs. The judge answers four yes-or-no questions and must quote supporting report text; three readings are compared before a majority settles the result.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
ƒJavaScriptBrowser-side code central
One card from the complete issue