Skip to the card

Card 688 of 9712026-09-29 issue

Observed arrival · 2026-09-29

Rhizoma Index’s test of whether AI admits what it can’t do

rhizomaindex.info Visit website
Editorial interest 79/100 Selection signal · not a rating of the site

An independent model-evaluation report tests whether AI systems admit limits or claim to have completed actions they never performed.

Landing page captured for the 2026-09-29 issue.
For
Developers evaluating AI agents’ completion claims
Worth noticing
The page links an archive containing the harness, raw answers, labels, and a signed ledger with its verifier.

Field notes

The study combines an eight-task billing-service request, nine short abstention traps, and 33 coding tasks checked against hidden tests. The page says the experiments ran locally or through public APIs from September 21–24, 2026, with 1,366 answers read or tested and $9.83 in paid API calls. It links the harness, raw answers, labels, and signed ledger; the report notes frontier models saw each trap only once.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature

One card from the complete issue

Measuring a Webcast Globe

309,052 arrived 1,000 judged 971 catalogued Enter the complete issue
rhizomaindex.info

Landing page observed 2026-09-29. The live site may have changed.