Observed arrival · 2026-09-30
AutoDataBench tests whether agents can write training tasks
A research benchmark evaluates agent-written executable tasks for validity, difficulty, and whether they elicit intended behavior.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI researchers and training-data builders
- Worth noticing
- The page reports Kimi K3’s score rising from 18.4 to 54.1 when its task-writing budget increases from 45 to 180 minutes.
Field notes
The described workflow starts with an existing task and the target model’s attempt, then asks an author agent to create a different task for the same suite. A gate rejects defects such as a verifier that never checks the result or a superficial rewrite; surviving tasks must also land within a six-attempt difficulty band. The results section says most tasks are either solved every time or not at all.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue