Observed arrival · 2026-09-08
The Measurement Gap
A 10,366-word research essay arguing that AI benchmark scores increasingly measure model-and-harness combinations rather than models alone.
Field notes
The project is organized as a long-form, sectioned essay rather than a continuously updated news feed. Its central measurement example compares the same stated model across different evaluation harnesses, while the appendices promise a worked metric example and a complete source list. The proposed DRIFT benchmark is described as measuring recovery from unannounced rule changes against a human baseline, but the page explicitly says that prototype has not yet been run. The author also dates predictions through 2028 so they can later be graded.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue