Skip to the card

Card 344 of 9772026-09-26 issue

Observed arrival · 2026-09-26

HalluWorld maps where language models still make things up

halluworld.com Visit website
Editorial interest 78/100 Selection signal · not a rating of the site

HalluWorld is a benchmark for measuring language-model hallucinations across grid worlds, chess, and realistic terminal tasks.

Landing page captured for the 2026-09-26 issue.
For
Researchers evaluating language-model reliability
Worth noticing
Its probes distinguish perceptual, causal, memory, and uncertainty errors across grid worlds, chess, and terminal tasks.

Field notes

The framework defines a hallucination as an observable claim that is false in a fully specified reference world. Its homepage says the environments allow controlled variation in world complexity, observability, temporal change, and source-conflict policy. Reported results distinguish near-solved perception from harder state tracking, forward simulation, and abstention; the page also cautions that extended thinking does not generally solve the causal probes.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

The Phone Disclosure Test

328,122 arrived 1,000 judged 977 catalogued Enter the complete issue
halluworld.com

Landing page observed 2026-09-26. The live site may have changed.