Observed arrival · 2026-09-17
Invigil, the benchmark a lab cannot train against
Invigil proposes machine-checked Lean 4 benchmarks for measuring whether frontier AI models can reason without simply recalling published test data.
Field notes
The proposed evaluation separates a reproducible public method from a private problem set that is never published and is refreshed quarterly from later research literature. Models receive theorem statements and return proof bodies; Invigil’s harness supplies the statement and Lean 4 checks correctness and axiom dependencies. Runs are described as pre-registered and hash-committed, with certification by Antefacts, a commonly owned entity whose relationship the site says is disclosed on its certification page.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue