Skip to the card

Card 578 of 9722026-09-25 issue

Observed arrival · 2026-09-25

Seven coding agents, one model, 84 public reports

openagentbench.com Visit website
Editorial interest 84/100 Selection signal · not a rating of the site

A public benchmark compares seven AI coding agents on the same 12 tasks using Kimi K3, with reports, code diffs, and hashes published for inspection.

Landing page captured for the 2026-09-25 issue.
For
AI coding-agent evaluators and developers
Worth noticing
The page says agent versions and inference settings were not fully logged, and excludes MiniMax from strict comparison pending a rerun.

Field notes

Task prompts are presented in their original Chinese, with per-task scoring bases archived separately. The page distinguishes a provisional five-part score from its frozen rule set, which weights acceptance, quality, verification, and safety; it says no composite should be issued while safety evidence is incomplete. It also records one run per task and leaves MiniMax outside strict comparison pending a rerun.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Vandalism in Demo Mode

315,813 arrived 1,000 judged 972 catalogued Enter the complete issue
openagentbench.com

Landing page observed 2026-09-25. The live site may have changed.