Observed arrival · 2026-09-23
Harnessbench replays your coding-agent tasks against new rules
A tool for comparing how changes to a coding assistant’s instructions affect the same saved coding tasks.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Teams tuning rules for AI coding assistants
- Worth noticing
- Its example compares the same saved tasks under old and new rules, including a rule that improves safety but leaves a task unfinished.
Field notes
The page walks through a saved-task comparison after a rule is added to a Claude Code instruction file. In its example, a task that requires moving eleven files becomes safer under the new rule but stalls after repeated permission requests, illustrating how the scoring weighs completion alongside safety. The example is explicitly labeled true-to-life, with names changed and times rounded; the homepage says runs happen on the user's machine.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
$PaidCommerce or pricing visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue