Observed arrival · 2026-09-02
LOL Bench Asks Whether Language Models Get the Joke
An open benchmark measuring how language models understand humor, produce jokes, and match human taste.
Field notes
The benchmark separates comprehension, production, and taste instead of collapsing humor into one model score. Explanations are checked against human-written joke mechanics by two family-disjoint AI judges, while generated jokes use 40 constrained premises and anonymous human pairwise voting. Scores require at least 10 judged explanations, and the page reports confidence intervals to limit small-sample overclaiming. The visible release is dataset v0.1.0, wave 0, with several ranking and calibration outputs still pending.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue