Skip to the card

Card 044 of 9982026-09-04 issue

Observed arrival · 2026-09-04

AppleBench asks whether agents can actually ship Apple apps

applebench.dev Observed source
Editorial interest 86/100 Selection signal · not a rating of the site

A benchmark tests whether AI agents can build, launch, operate, and leave behind a working Apple-platform app.

Landing page captured for the 2026-09-04 issue.

Field notes

AppleBench evaluates an agent’s work after the agent exits, using a fresh build and tests against the workspace it leaves behind. Its 134 tasks span eight categories, including project configuration, simulator interaction, runtime defects, visual layout, and Apple frameworks; 45 tasks are graded by running the app. The site says reference fixes and no-change failures are checked with real xcodebuild before inclusion, while reported token and cost data remain null when the CLI provides none.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

An Octave Below the Spacecraft

377,806 arrived 1,000 judged 998 catalogued Enter the complete issue
applebench.dev

Landing page observed 2026-09-04. The live site may have changed.