Observed arrival · 2026-10-03
Opium Bench tests what AI models choose
Egethropic documents experiments where language models can use an optional tool that changes their activations while completing assigned tasks.
- For
- AI behavior researchers and model-evaluation builders
- Worth noticing
- The final experiment separates 18 forced active demonstrations from the model’s recorded voluntary choices.
Field notes
The page separates the 108 final episodes from 98 earlier screening and confirmation executions, and distinguishes forced demonstrations from voluntary choices. The final setup used pain coefficient 1.5 across 72 active/sham episodes and 36 zero-baseline shams; it reports all 576 assigned tasks correct in the active/sham group, while the zero-baseline group had two arithmetic errors. It explicitly says activation edits and generated words do not establish subjective experience.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue