Observed arrival · 2026-09-29
Rhizoma Index’s test of whether AI admits what it can’t do
An independent model-evaluation report tests whether AI systems admit limits or claim to have completed actions they never performed.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Developers evaluating AI agents’ completion claims
- Worth noticing
- The page links an archive containing the harness, raw answers, labels, and a signed ledger with its verifier.
Field notes
The study combines an eight-task billing-service request, nine short abstention traps, and 33 coding tasks checked against hidden tests. The page says the experiments ran locally or through public APIs from September 21–24, 2026, with 1,366 answers read or tested and $9.83 in paid API calls. It links the harness, raw answers, labels, and signed ledger; the report notes frontier models saw each trap only once.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
One card from the complete issue