Observed arrival · 2026-09-16
Model Field Guide, for looking past the leaderboard
A practical reference for evaluating AI systems by task fit, evidence, limits, and the cost of results.
Field notes
The guide frames evaluation as a sequence of task definition, evidence checking, and boundary testing rather than as a single model-ranking exercise. Its diagnostic material names separate failure layers—including retrieval, OCR, schemas, tools, permissions, and interfaces—while the accepted-result calculator keeps calculations in the browser using costs, review time, retries, and acceptance rate. A historical Stanford CRFM diagram is offered as context, with an explicit warning that it is not a benchmark result.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue