Observed arrival · 2026-09-25
Seven coding agents, one model, 84 public reports
A public benchmark compares seven AI coding agents on the same 12 tasks using Kimi K3, with reports, code diffs, and hashes published for inspection.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI coding-agent evaluators and developers
- Worth noticing
- The page says agent versions and inference settings were not fully logged, and excludes MiniMax from strict comparison pending a rerun.
Field notes
Task prompts are presented in their original Chinese, with per-task scoring bases archived separately. The page distinguishes a provisional five-part score from its frozen rule set, which weights acceptance, quality, verification, and safety; it says no composite should be issued while safety evidence is incomplete. It also records one run per task and leaves MiniMax outside strict comparison pending a rerun.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue