Observed arrival · 2026-09-05
ReasonEval grades the reasoning, not just the answer
An open benchmark suite for evaluating AI models on math, code, logic, planning, calibration, and agentic tasks.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
The site divides evaluation into six suites, including formal-verified mathematics, sandboxed coding, long-horizon planning, calibration, and replayable agent tasks. Its displayed composite score weights correctness, process, calibration, and robustness rather than treating accuracy as the sole measure. The command-line example accepts an OpenAI-compatible endpoint and writes a trace-graded HTML report with per-suite results and uncertainty ranges.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue