Skip to the card

Card 523 of 9982026-09-08 issue

Observed arrival · 2026-09-08

The Measurement Gap

measurementgap.com Observed source
Editorial interest 84/100 Selection signal · not a rating of the site

A 10,366-word research essay arguing that AI benchmark scores increasingly measure model-and-harness combinations rather than models alone.

Landing page captured for the 2026-09-08 issue.

Field notes

The project is organized as a long-form, sectioned essay rather than a continuously updated news feed. Its central measurement example compares the same stated model across different evaluation harnesses, while the appendices promise a worked metric example and a complete source list. The proposed DRIFT benchmark is described as measuring recovery from unannounced rule changes against a human baseline, but the page explicitly says that prototype has not yet been run. The author also dates predictions through 2028 so they can later be graded.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

1985 at 160× Speed

370,686 arrived 1,000 judged 998 catalogued Enter the complete issue
measurementgap.com

Landing page observed 2026-09-08. The live site may have changed.