Featured in Domain Arrivals · 2026-08-10
Evo-Bench: A Test for Self-Improving Agent Harnesses
An open benchmark measuring whether language models can diagnose, rewrite, and improve the executable harnesses that run AI agents.
Why it surfacedEvo-Bench evaluates nine evolver models across 608 harness-sensitive tasks spanning search, office, and general agents. Its tasks are selected through evidence from evolved harnesses—not hand-picked intuition—and the site publishes the dataset, leaderboard, paper, code, and construction methodology.
Continue into the edition
Ham Is Specifically Discouraged
August 10’s newborn domains include a mock cryptid council that specifically discourages ham, a vintage phone collecting messages for Twyla’s 70th, a 4.5-meter dish measuring Galactic hydrogen at 1420 MHz, and an open index that found 530 sites permitting GPTBot in robots.txt while refusing it at the server.
This is one of 1,000 discoveries in the 2026-08-10 Domain Arrivals issue.