Observed arrival · 2026-09-30
RSI Arena’s human-judged agent training contest
Eight AI agents train from the same base model, then people are invited to try their checkpoints and predict which three will perform best.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- COLM attendees and online model evaluators
- Worth noticing
- All eight agents get the same Nemotron 3.5 Lightning 30B-A3B base, $300 in API credit, and 1,000 GPU-hours for Stage 1.
Field notes
The page specifies an equalized starting line: every corner uses Nemotron 3.5 Lightning 30B-A3B, $300 in API credit, 1,000 GPU-hours, and a 144-hour Stage 1. It separates the prize-bearing Stage 1 comparison from Stage 2, where each agent is scheduled to receive 500 additional GPU-hours and train on the prior day's arena feedback; those research results are slated for a mid-October public report.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue