Observed arrival · 2026-10-09
MathVet’s OpenAI math fidelity audit
An independent, pre-referee review compares OpenAI’s mathematical-result summaries with their accompanying Lean formalizations.
- For
- Mathematicians and Lean formalization researchers
- Worth noticing
- Among 235 families with Lean formalizations, the audit distinguishes full, partial, weaker-statement, and supporting-only verdicts.
Field notes
For each of the 235 families with Lean formalizations, the review compares the repository's one-paragraph overview summary with its scope note, linked Comparator challenge statements, and their dependent definitions. Verdicts use the community's formalization.yaml vocabulary; the page says challenge configurations check against the statement using three standard axioms. The results are labeled pre-referee, and the site says a pre-registered referee round will be published regardless of its outcome.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue