Observed arrival · 2026-09-12
Pelican Arena: Blind-Testing Models on a Bicycle-Drawing Challenge
A Chinese-language arena compares AI-generated SVG drawings of a pelican riding a bicycle through anonymous head-to-head voting.
Field notes
The benchmark uses a deliberately awkward SVG assignment—drawing a pelican riding a bicycle—to test whether models can construct spatially coherent vector code rather than reproduce an obvious existing snippet. The visible interface pairs two concealed submissions, supports keyboard voting and ties, and says identities appear after the vote. It also names a dynamic Elo ladder, an eight-model collection, a gallery, and API checks for latency, throughput, and alleged chain-of-thought impersonation.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue