Observed arrival · 2026-09-27
WTF Evals turns LLM failures into test plans
A plain-English guide helps you turn a specific LLM failure into a small, testable evaluation and a decision rule.
- For
- Teams evaluating LLM features before release
- Worth noticing
- Its example blocks release after one of 12 tasks weakens authorization, despite a 10/12 pass rate.
Field notes
The example eval checks ordinary backend tasks involving authenticated endpoints, with original tests, hidden authorization checks, and a protected-file diff as evidence. Its fill-in sketch asks users to name a decision, an observable failure, relevant tasks, a way to observe the failure, and an action threshold; the guide also cautions that small case sets are for learning, not proof of universal quality. It distinguishes direct software checks from human review and LLM judges, which it says should be calibrated against independent human labels.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue