Observed arrival · 2026-10-05
FloppyBench puts AI models inside classic games
A benchmark project runs AI models through unmodified classic games, using each game’s own rules to judge their actions.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- AI benchmark builders and classic-game enthusiasts
- Worth noticing
- The harness reads game state from emulator memory and logs the briefing, model reasoning, tool calls, and game responses.
Field notes
The project separates solo benchmarks from competitions in which models share a game world; it says competition results do not affect benchmark rankings. Its stated logging covers the briefing, model reasoning, tool calls, and game responses, while the harness reads state from the running emulator. The first competition is still listed as “to be announced,” and the homepage extract does not show current scores.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
⊠LoginAccess appeared gated
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue