Observed arrival · 2026-08-26
Fluency Bench: an AI test that tries to catch the bluff
A 25-minute, three-round work sample that tests whether someone can brief an AI model, detect fabricated claims, and build a process that survives unseen cases.
Field notes
The assessment divides AI work into briefing, verification, and system-building rather than treating prompting as a single skill. Its published scoring includes six dimensions, a cost ceiling of $0.15 per item, and a visible overfit_delta comparing performance on seen and unseen cases. A sample report also shows adversarial results and category-level scores. The homepage says access codes travel by invite, while the rubric and sample evidence remain publicly inspectable.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue