Observed arrival · 2026-09-18
Business Bench Measures Whether Agents Actually Finish the Work
An executable benchmark for testing AI agents on business tasks across documents, spreadsheets, bookkeeping, reporting, extraction, and drafting.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI evaluation researchers and business-automation builders
- Worth noticing
- Each task runs three fresh-workspace repetitions and publishes raw and frozen verdicts bound to a hashed scorer.
Field notes
Business Bench evaluates generated business tasks through exact checks applied to the files an agent produces, rather than rubric scores or an LLM judge. It repeats each task three times in fresh workspaces and tracks both conjunctive passing and partial check performance, alongside wall time, token work, and estimated cost. The homepage links to methods, reproduction materials, an audit, a paper, and a public repository.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue