Observed arrival · 2026-09-04
NextChapterBench Tests Whether Models Can Continue a Novel
A literary-continuation benchmark ranks language models on how well they write a deliberately divergent next chapter.
Field notes
The benchmark supplies each model with a public-domain novel chapter plus a deliberately divergent brief, then compares the resulting continuations across compliance and craft. Its standings preserve operational detail often omitted from model comparisons: reasoning-effort setting, CLI route, output word count, per-run cost, and uncertainty margins. The visible table covers 13 models, 10 items, and three LLM judges, with links leading to model detail pages and the underlying chapters.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue