Observed arrival · 2026-09-21
RRSI: Regularized Recursive Self-Improvement
A research project on evolving AI agent harnesses without letting them overfit the benchmark used to score them.
- For
- AI researchers and agent-harness builders
- Worth noticing
- The reported held-out gains span six benchmarks that the harness never optimizes.
Field notes
The project frames agent-harness evolution as a problem of search regularization rather than restricting the harness components themselves. Its visible method includes an annealed edit budget, an evidence ledger, a leakage critic, a noise-adjusted floor, and a cost rule for additional inference tokens. The results section distinguishes three benchmarks used for evolution from six held-out benchmarks, covering coding, agentic workspace, and engineering-design tasks.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue