Skip to the card

Card 476 of 9992026-09-02 issue

Observed arrival · 2026-09-02

LOL Bench Asks Whether Language Models Get the Joke

lolbench.lol Observed source
Editorial interest 86/100 Selection signal · not a rating of the site

An open benchmark measuring how language models understand humor, produce jokes, and match human taste.

Landing page captured for the 2026-09-02 issue.

Field notes

The benchmark separates comprehension, production, and taste instead of collapsing humor into one model score. Explanations are checked against human-written joke mechanics by two family-disjoint AI judges, while generated jokes use 40 constrained premises and anonymous human pairwise voting. Scores require at least 10 judged explanations, and the page reports confidence intervals to limit small-sample overclaiming. The visible release is dataset v0.1.0, wave 0, with several ranking and calibration outputs still pending.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

A Commons Made of Smoke

312,505 arrived 1,000 judged 999 catalogued Enter the complete issue
lolbench.lol

Landing page observed 2026-09-02. The live site may have changed.