Observed arrival · 2026-08-28
WhichBuildWon Puts AI Coding Agents Through the Same Tests
A comparison site gives AI coding agents identical briefs, publishes the resulting applications, and checks what actually worked.
Field notes
The site standardizes comparisons by using a clean starting commit, one prompt, one fresh non-resumed CLI session, and the same deterministic checks for each contestant. Its homepage links to full captures and reports concrete outcomes, including a Relay landing-page run in which Codex passed all 20 checks while Claude Code missed two. Another experiment holds the model and task constant while changing only the presence of a root CLAUDE.md file.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue