Observed arrival · 2026-09-03
Integration Bench tests whether AI agents can survive real API integrations
A benchmark gives AI coding agents integration tasks against 15 synthetic vendors, then grades the resulting behavior from vendor-side request logs.
Field notes
The benchmark runs agents against controlled vendors and judges the resulting HTTP behavior from request logs, rather than trusting the agent's source code alone. The public release covers 50 tasks across 15 synthetic vendors and 17 models, with full tool calls, vendor requests, graded checks, and diffs for 850 attempts. The site identifies the release as rev 01 and preliminary, while describing a held-out set and additional evidence bundles as early-access material.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue