Skip to the card

Card 457 of 10002026-09-05 issue

Observed arrival · 2026-09-05

LLM Proof-Grading Study

llmproofbiasstudy.website Observed source
Editorial interest 84/100 Selection signal · not a rating of the site

A public research record examining how names and gender cues affect language-model grading of mathematical proofs.

Landing page captured for the 2026-09-05 issue.

Field notes

The study compares four language models on ten proof attempts, using ten named male and ten named female students per model alongside anonymous-baseline conditions. Scores are normalized against each problem’s anonymous mean and standard deviation, with separate model-level, problem-level, and student-level views. The site also preserves raw LaTeX, rendered proofs, result workbooks, and named-student and anonymous CSV files. One result, Grok Problem 4, is excluded from normalized analysis because its anonymous standard deviation is zero.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

The Receiver Chooses a Memory

351,339 arrived 1,000 judged 1000 catalogued Enter the complete issue
llmproofbiasstudy.website

Landing page observed 2026-09-05. The live site may have changed.