Observed arrival · 2026-09-10
benchgraph turns AI benchmarks into a readable graph
An open catalogue maps what AI benchmarks measure, who publishes them, which models they cover, and how current their scores are.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
The graph treats a benchmark as more than a score: its visible schema includes the task, dataset, licence, publisher, lineage, successors, model coverage, and a freshness date. The homepage’s SWE-bench Verified example records 500 tasks from 12 Python repositories and shows 112 models with scores. The project says its current preview uses ModelSpec cards, with dedicated benchmark pages to be populated by a daily research process.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue