Skip to the card

Card 602 of 9742026-10-01 issue

Observed arrival · 2026-10-01

Plumbline Grader puts code-RL graders on trial

plumblinegrader.com Visit website
Editorial interest 82/100 Selection signal · not a rating of the site

Plumbline Grader audits the tests used to score code-reinforcement-learning tasks, looking for wrong patches that still earn full reward.

Landing page captured for the 2026-10-01 issue.
For
Teams building code-RL environments and task graders
Worth noticing
The homepage distinguishes 78 strict-scope failures from 143 total findings across 191 tested tasks in its 200-task sample.

Field notes

The audit separates failures required by an issue as written (“strict”) from broader test gaps, and says it reports those categories separately. Wrong patches come from two sources: incomplete pieces of the reference fix and model-written plausible bugs or partial fixes. The site says findings are tied to run logs, with a sample rerun in fresh containers; its described deliverables include per-task verdicts and test patches checked against known wrong patches.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Pink Dolphin Slippers

676,985 arrived 1,000 judged 974 catalogued Enter the complete issue
plumblinegrader.com

Landing page observed 2026-10-01. The live site may have changed.