Observed arrival · 2026-09-15
ExplorationBench Makes AI Explore Alien Rules
An open benchmark tests whether AI systems can discover hidden semantics in executable worlds rather than rely on memorized knowledge.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
ExplorationBench separates discovering a rule from merely recalling a familiar domain. AlienCode tests programs against private inputs through an interpreter, while AlienLogic checks Fitch-style proofs and includes 21 unprovable goals where refusal can be correct. Its protocol compares four rounds of probing across twenty system–sandbox pairs, with each leaderboard cell averaged over three runs.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue