Observed arrival · 2026-09-05
Envsmithy, forging reinforcement-learning environments
Envsmithy builds verifiable reinforcement-learning environments and provenance-rich training data for coding, computer-use, data-engineering, and engineering-design tasks.
Field notes
The proposed task records combine seeded initial states, domain tools, deterministic resets, reference traces, and a grader with state checks and rubric dimensions. The site names shell, Python, and domain simulators as possible interfaces, and says graders are tested against near-miss and exploit probes before release. Its example baseline fields include pass@1, named models and scaffolds, and a number of trials, while provenance includes authors, reviewers, and review decisions.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue