Skip to the card

Card 675 of 10002026-09-05 issue

Observed arrival · 2026-09-05

ReasonEval grades the reasoning, not just the answer

reasoneval.com Observed source
Editorial interest 82/100 Selection signal · not a rating of the site

An open benchmark suite for evaluating AI models on math, code, logic, planning, calibration, and agentic tasks.

Landing page captured for the 2026-09-05 issue.

Field notes

The site divides evaluation into six suites, including formal-verified mathematics, sandboxed coding, long-horizon planning, calibration, and replayable agent tasks. Its displayed composite score weights correctness, process, calibration, and robustness rather than treating accuracy as the sole measure. The command-line example accepts an OpenAI-compatible endpoint and writes a trace-graded HTML report with per-suite results and uncertainty ranges.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

The Receiver Chooses a Memory

351,339 arrived 1,000 judged 1000 catalogued Enter the complete issue
reasoneval.com

Landing page observed 2026-09-05. The live site may have changed.