Observed arrival · 2026-09-16
AI Scientist Bench: A Ledger for Scientific AI Skills
An open benchmark comparing AI systems on scientific tasks including claim checking, paper selection, information extraction, and data analysis.
Field notes
The benchmark separates scientific abilities into claim checking, paper selection, information extraction, and data analysis, using task-specific measures rather than one aggregate score. Its records include sample sizes, model-call cost, elapsed time, and treatment of failures; one extraction test even reports a post-hoc wrapper rescore. The analysis reproduction is deliberately narrow: one supplied-code structural-equation case, three setups, and checks using both original and changed inputs.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue