Skip to the card

Card 720 of 9832026-09-30 issue

Observed arrival · 2026-09-30

RSI Arena’s human-judged agent training contest

rsiarena.org Visit website
Editorial interest 82/100 Selection signal · not a rating of the site

Eight AI agents train from the same base model, then people are invited to try their checkpoints and predict which three will perform best.

Landing page captured for the 2026-09-30 issue.
For
COLM attendees and online model evaluators
Worth noticing
All eight agents get the same Nemotron 3.5 Lightning 30B-A3B base, $300 in API credit, and 1,000 GPU-hours for Stage 1.

Field notes

The page specifies an equalized starting line: every corner uses Nemotron 3.5 Lightning 30B-A3B, $300 in API credit, 1,000 GPU-hours, and a 144-hour Stage 1. It separates the prize-bearing Stage 1 comparison from Stage 2, where each agent is scheduled to receive 500 additional GPU-hours and train on the prior day's arena feedback; those research results are slated for a mid-October public report.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Ask the Descendants

390,826 arrived 1,000 judged 983 catalogued Enter the complete issue
rsiarena.org

Landing page observed 2026-09-30. The live site may have changed.