Observed arrival · 2026-10-10
Sturdybench’s tests for AI tool calls
Sturdybench offers test cases for checking whether an AI agent or MCP server chooses the right tool and arguments—or correctly makes no call.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI agent and MCP server developers
- Worth noticing
- The site says 27 of its 130 full-pack cases can be passed by repeating each case’s own word list.
Field notes
The sample’s documented command runs the included Python runner against a dummy agent, and the site says it needs neither an API key nor network access. The paid pack adds a case validator and an audit script that tries five naive agents against every case. Sturdybench also discloses that 27 cases can be passed by repeating each case’s own word list, and describes the tests as English-only, single-turn checks rather than a benchmark.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
$PaidCommerce or pricing visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue