Observed arrival · 2026-10-10
CoBe: Testing Counterfactual Editing in Language Models
CoBe is a benchmark for testing whether language models can revise a story after a hypothetical change while preserving details that change should not affect.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Researchers evaluating language-model reasoning
- Worth noticing
- The site contrasts Pearl's causal-ladder Rung 3 counterfactuals with Rung 1 associative rewrites using paired examples.
Field notes
The task is presented as a three-part operation: preserve unaffected background facts, enact the hypothetical intervention, and update downstream consequences. One example contrasts a counterfactual about Kennedy surviving with an associational rewrite that invents another shooter. The site also gives a medical-record scenario to illustrate how changing symptoms instead of treatment could distort a downstream training case.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue