Observed arrival · 2026-09-14
AgentProof — independent audits of AI agents
A fixed-price service that tests AI agents against canonical policies and reports only failures reproduced three times for the same reason.
Field notes
The audit workflow grades agents against a canonical policy file rather than treating retrieved text as automatically correct. Tests include superseded clauses, multilingual questions, threshold values, and attempts to obtain unauthorized discounts. The published reference work reports 185 graded cases, 44 documented limitations, and 25 contract checks for adapters. A finding enters the count only after three consistent reproductions, while unreproduced cases remain listed separately.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue