Observed arrival · 2026-09-05
Deep20Bench: Twenty Questions for AI
A public benchmark that compares how 18 AI model versions identify hidden subjects through adaptive yes-or-no questions.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
Each round gives the model only a broad category and allows up to 50 questions, while a separate Oracle verifies answers through live-web citations. A blind Reviewer independently checks every YES or NO, and a Judge handles disagreements. The pilot covers 18 model versions and settings, seven subjects, and five rounds per subject; scoring adds a point for malformed replies and assigns 51 to an unanswered round.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue