Skip to the card

Card 013 of 9802026-09-27 issue

Observed arrival · 2026-09-27

AgenticBench puts coding agents through a public audit

agenticbench.org Visit website
Editorial interest 86/100 Selection signal · not a rating of the site

A versioned benchmark tests what AI coding agents record, send off a machine, and do when no person is watching.

Landing page captured for the 2026-09-27 issue.
For
People evaluating AI coding agents for real work
Worth noticing
A correction audit moved 21 previously published passes to “Not tested by us”; no published fail changed.

Field notes

The scoring separates an actual failure from a test the lab could not complete, and says the latter never counts against an agent. Results are version-specific rather than presented as a certification. The page also records changes and corrections: after reviewing its capture pipeline, the lab withdrew 21 passes that its evidence did not fully support, while reporting that no published fail changed.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Nostalgia for an Imaginary Console

276,168 arrived 1,000 judged 980 catalogued Enter the complete issue
agenticbench.org

Landing page observed 2026-09-27. The live site may have changed.