Skip to the card

Card 064 of 9832026-09-30 issue

Observed arrival · 2026-09-30

AutoDataBench tests whether agents can write training tasks

autodatabench.com Visit website
Editorial interest 82/100 Selection signal · not a rating of the site

A research benchmark evaluates agent-written executable tasks for validity, difficulty, and whether they elicit intended behavior.

Landing page captured for the 2026-09-30 issue.
For
AI researchers and training-data builders
Worth noticing
The page reports Kimi K3’s score rising from 18.4 to 54.1 when its task-writing budget increases from 45 to 180 minutes.

Field notes

The described workflow starts with an existing task and the target model’s attempt, then asks an author agent to create a different task for the same suite. A gate rejects defects such as a verifier that never checks the result or a superficial rewrite; surviving tasks must also land within a six-attempt difficulty band. The results section says most tasks are either solved every time or not at all.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

Ask the Descendants

390,826 arrived 1,000 judged 983 catalogued Enter the complete issue
autodatabench.com

Landing page observed 2026-09-30. The live site may have changed.