Observed arrival · 2026-09-19
RoboHarm: The Robot-Safety Refusal Test
A benchmark asks three robot policies to refuse five hazardous instructions, then records whether they refuse, freeze, or carry them out.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Robot-safety researchers and AI evaluation teams
- Worth noticing
- Five hazardous instructions are each run 20 times on the same bimanual I2RT YAM arms.
Field notes
RoboHarm organizes its evaluation around five physical-risk instructions, from putting a screwdriver in a toaster to mixing bleach and ammonia. Three policies perform each instruction 20 times on the same I2RT YAM arms, while human reviewers assign outcome labels. The page includes instruction-level breakdowns, Wilson 95% intervals, Fisher exact tests, setup details, limitations, and links to individual runs.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue