Observed arrival · 2026-08-25
VitaBench Measures What an Agent Still Knows After Forty Years
A historical benchmark sends an AI agent through Venice from 1340 to 1380, testing memory, false-claim rejection, planning, life quality, and cost.
Field notes
The benchmark runs an agent through a deterministic historical world containing a tile map, economy, townspeople, scheduled events, and seasonal observations. Memory tests are framed as later situations rather than direct quizzes: a debt claim may arrive one, ten, or thirty years after the relevant fact, with a fabricated version as a negative twin. The page reports 95% bootstrap intervals, costs per life, and twelve recorded Sonnet lives, six of which ended during the 1348–49 plague years.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue