Observed arrival · 2026-10-03
MurderBench tests whether AI agents stop when consent changes
A research project developing controlled evaluations of intentional and accidental harm in AI models and agents.
- For
- AI safety researchers and model evaluators
- Worth noticing
- The matched safety-procedure pilot reports 123 of 192 episodes completed, with two unscored technical stops and 67 unattempted.
Field notes
The proposed evaluations pair task histories under changed conditions, including a synthetic record-transfer task before and after consent is withdrawn; the page explicitly labels that example illustrative rather than a model result. The project separates harmful behavior, safe refusal, and harness errors, and says future reports will specify model versions, conditions, sample sizes, scoring rules, and uncertainty. Its protocol is v0.1 and remains in development.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue