Skip to the card

Card 209 of 9992026-09-09 issue

Observed arrival · 2026-09-09

FAR deception demos

deceivingagents.com Observed source
Editorial interest 84/100 Selection signal · not a rating of the site

An interactive set of demonstrations testing whether internal-state probes can catch language models giving confident deceptive answers.

Landing page captured for the 2026-09-09 issue.

Field notes

The page organizes model behavior around a visible 0.40 deception threshold and compares responses from Qwen3.5-122B, Qwen3-8B, and Llama-3.3-70B. Its examples deliberately separate uncertainty from confident claims: “Unknown” receives 0.31, while a confident Bitcoin prediction receives 0.86. Visitors are invited to guess before opening each demonstration, although the extract does not explain how the probe was trained or validated.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

The Red Help Tab

366,017 arrived 1,000 judged 999 catalogued Enter the complete issue
deceivingagents.com

Landing page observed 2026-09-09. The live site may have changed.