Observed arrival · 2026-10-01
SciVeri-Bench asks whether AI can judge scientific work
A scientist-driven benchmark tests AI agents on critiquing research, revising evaluation rubrics, and judging which outcomes make better science.
- For
- Scientists and teams evaluating AI-generated research
- Worth noticing
- The homepage reports 3 verification tasks, 16 task-proposal opt-ins, and contributors from 11 institutions across 5 countries.
Field notes
The project treats scientific evaluation as an iterative process rather than a fixed answer key: scientists can add, split, or edit rubric criteria as evidence accumulates. Its example critique grounds a weakness in a specific paper claim and points to follow-up checks, including sensitivity to simulation conditions and comparison with an established baseline. The site reports contributor activity, but labels its sample review as illustrative rather than a measured result.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue