Skip to the card

2026-10-10 issue

Observed arrival · 2026-10-10

System-One Control Bench puts AI decisions through 100 grid puzzles

socb.dev Visit website
Editorial interest 79/100 Selection signal · not a rating of the site

A public benchmark compares decision models, chat models, and simple programmed strategies as they guide an agent through grid puzzles.

Landing page captured for the 2026-10-10 issue.
For
AI benchmark researchers and model evaluators
Worth noticing
The report says Jev chooses optimal moves in about 90% of separate test positions but finishes only 57 of 100 full games.

Field notes

Each model runs the same 100 puzzles under map-only and full-context conditions, with wins, mean progress, and API cost reported across 200 games. A solver supplies shortest-route references, letting the benchmark distinguish locally optimal moves from successful end-to-end play; saved game recordings can be replayed without making new model requests.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Voynich on Demand

385,880 arrived 1,000 judged 985 catalogued Enter the complete issue
socb.dev

Landing page observed 2026-10-10. The live site may have changed.