Observed arrival · 2026-09-09
FAR deception demos
An interactive set of demonstrations testing whether internal-state probes can catch language models giving confident deceptive answers.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
The page organizes model behavior around a visible 0.40 deception threshold and compares responses from Qwen3.5-122B, Qwen3-8B, and Llama-3.3-70B. Its examples deliberately separate uncertainty from confident claims: “Unknown” receives 0.31, while a confident Bitcoin prediction receives 0.86. Visitors are invited to guess before opening each demonstration, although the extract does not explain how the probe was trained or validated.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue