Observed arrival · 2026-09-04
AppleBench asks whether agents can actually ship Apple apps
A benchmark tests whether AI agents can build, launch, operate, and leave behind a working Apple-platform app.
Field notes
AppleBench evaluates an agent’s work after the agent exits, using a fresh build and tests against the workspace it leaves behind. Its 134 tasks span eight categories, including project configuration, simulator interaction, runtime defects, visual layout, and Apple frameworks; 45 tasks are graded by running the app. The site says reference fixes and no-change failures are checked with real xcodebuild before inclusion, while reported token and cost data remain null when the CLI provides none.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue