Observed arrival · 2026-09-07
Copperbench tests AI agents on actual circuit boards
Copperbench benchmarks language-model agents that edit KiCad hardware designs, checking their work with kicad-cli and publishing pass rates alongside cost.
Field notes
Copperbench separates benchmark validity from model performance by checking edits offline with kicad-cli and fixed assertions. Its records are only comparable when suite version, task manifest hash, fixture hash, and kicad-cli major version match. The site currently reports sixteen harness records that separate no-op behavior from correct edits, while its simple and medium leaderboard tiers have not yet recorded model runs and its hard tier contains no tasks.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue