Observed arrival · 2026-10-10
System-One Control Bench puts AI decisions through 100 grid puzzles
A public benchmark compares decision models, chat models, and simple programmed strategies as they guide an agent through grid puzzles.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI benchmark researchers and model evaluators
- Worth noticing
- The report says Jev chooses optimal moves in about 90% of separate test positions but finishes only 57 of 100 full games.
Field notes
Each model runs the same 100 puzzles under map-only and full-context conditions, with wins, mean progress, and API cost reported across 200 games. A solver supplies shortest-route references, letting the benchmark distinguish locally optimal moves from successful end-to-end play; saved game recordings can be replayed without making new model requests.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue