Observed arrival · 2026-09-29
TensorTest puts video-editing agents to a blunt test
TensorTest offers AI labs private evaluations, RL environments, and training data based on real creative work; its first benchmark measures video-editing agents.
- For
- AI labs evaluating creative-work agents
- Worth noticing
- TimelineBench's quality test is calibrated against blind judgments from 43 professional editors.
Field notes
TimelineBench gives agents raw footage, audio, and a brief, then checks whether the finished cut meets delivery, content, and brief requirements. Its quality test is calibrated against blind judgments from 43 professional editors, and the homepage links to individual runs and a methods write-up. TensorTest also describes private checkpoint evaluations, team-scoped RL environments, and commissioned training and evaluation data, with terms scoped per engagement.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue