Observed arrival · 2026-09-09
Agentic Robotics Benchmark: an exam for autonomous machines
A robotics evaluation platform that runs autonomous agents through sandboxed tasks and scores their resulting controllers, plans, designs, or code on hidden seeds.
Field notes
The benchmark separates agent work from evaluation: the ALE engine blocks networking, freezes /home/user/submission, and runs verification only after the agent stage ends. Scores are normalized against a reference implementation's measured anchor, capped, and aggregated with skipped tasks counting as zero. The homepage currently shows drone_hover as a UAV Control task with a three-hour agent budget and mean episode reward as its metric.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue