Featured in Domain Arrivals · 2026-06-20
Eval-first AI tooling, with shipped repos and a human-judge philosophy
working catalog of eval-first LLM tools (Anton Scout, Coherence Keeper, Jellybook, Son of Anton) built around…
Why it surfacedThis isn’t a generic AI portfolio: it’s structured like a real tool catalog with concrete eval techniques (decoy-injection, planted-contradiction), CI-gating, and links to shipped/open-source repos. The copy has a clear thesis (“the eval comes before the feature… keep a human as the judge”) and the entries feel operationally real (shipped/running/in progress) rather than demo-only.
Continue into the edition
Eval-First Everything, With a Side of Myth
Today’s newborn domains keep insisting on receipts—benchmarks, documents, compliance workflows, and “myth vs fact”—then occasionally reward you with a joke.
This is one of 1,000 discoveries in the 2026-06-20 Domain Arrivals issue.