Observed arrival · 2026-09-16
BonuslyBench puts 103 models through revenue-operations work
An inspectable benchmark tests language models on 40 real-world GTM tasks, including forecasts, CRM audits, renewal-risk calls, and executive briefs.
Field notes
BonuslyBench organizes its evaluation as an auditable package rather than a single leaderboard. Its 40 tasks cover forecasts, CRM audits, renewal-risk decisions, and executive briefs, while the reviewer exposes each model’s unedited response alongside the expected answer and individual checks. The site also supplies a four-sheet results workbook, seven analysis charts, model archetypes, and per-task deltas against claude-sonnet-5. It cautions that comparisons apply only to this 103-model, one-run sample.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue