Observed arrival · 2026-09-05
LLM Proof-Grading Study
A public research record examining how names and gender cues affect language-model grading of mathematical proofs.
Field notes
The study compares four language models on ten proof attempts, using ten named male and ten named female students per model alongside anonymous-baseline conditions. Scores are normalized against each problem’s anonymous mean and standard deviation, with separate model-level, problem-level, and student-level views. The site also preserves raw LaTeX, rendered proofs, result workbooks, and named-student and anonymous CSV files. One result, Grok Problem 4, is excluded from normalized analysis because its anonymous standard deviation is zero.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue