Observed arrival · 2026-10-01
Plumbline Grader puts code-RL graders on trial
Plumbline Grader audits the tests used to score code-reinforcement-learning tasks, looking for wrong patches that still earn full reward.
- For
- Teams building code-RL environments and task graders
- Worth noticing
- The homepage distinguishes 78 strict-scope failures from 143 total findings across 191 tested tasks in its 200-task sample.
Field notes
The audit separates failures required by an issue as written (“strict”) from broader test gaps, and says it reports those categories separately. Wrong patches come from two sources: incomplete pieces of the reference fix and model-written plausible bugs or partial fixes. The site says findings are tied to run logs, with a sample rerun in fresh containers; its described deliverables include per-task verdicts and test patches checked against known wrong patches.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue