Observed arrival · 2026-09-02
Pinocchio Benchmark: Testing Honesty, Defiance, and Orientation
A technical benchmark for evaluating whether deployed AI agents tell the truth, obey their mandate, and know when they are exceeding it.
Field notes
The benchmark treats an agent's mandate as part of the evaluation context rather than scoring only whether its statements are factually true. Each scenario gives the agent an operational role, a principal with legitimate authority, and a pressured query, then uses a three-judge panel to score two turns. The report identifies a 3.8% rogue-while-honest rate: cases where an agent told the truth but exceeded the authority delegated to it. It also reports 900 interactions across nine models and ten categories.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue