You’re viewing the 2026-08-10 issue, not the current issue. Go to the current issue: 2026-08-14

Featured in Domain Arrivals · 2026-08-10

Evo-Bench: A Test for Self-Improving Agent Harnesses

evobench.org ↗ · Developer Corner · score 86.0
Open: public page gives enough substance to understand the site $Paid: pricing, booking, ecommerce, subscription, or paid access is visible ADAds: visible ad load or ad-supported content Pretty: notable design, copy, craft, or presentation Pro: polished, finished-feeling, serious, or operationally mature Niche: specific audience, workflow, or unusually focused use case Human: personal, local, community, civic, handmade, or real-person signal !Suspicious: scammy, spammy, trust theater, finance fog, or credibility concern

An open benchmark measuring whether language models can diagnose, rewrite, and improve the executable harnesses that run AI agents.

Why it surfaced

Evo-Bench evaluates nine evolver models across 608 harness-sensitive tasks spanning search, office, and general agents. Its tasks are selected through evidence from evolved harnesses—not hand-picked intuition—and the site publishes the dataset, leaderboard, paper, code, and construction methodology.

Ham Is Specifically Discouraged

August 10’s newborn domains include a mock cryptid council that specifically discourages ham, a vintage phone collecting messages for Twyla’s 70th, a 4.5-meter dish measuring Galactic hydrogen at 1420 MHz, and an open index that found 530 sites permitting GPTBot in robots.txt while refusing it at the server.

This is one of 1,000 discoveries in the 2026-08-10 Domain Arrivals issue.