Observed arrival · 2026-09-05
Bufo Benchmark: Who Makes the Best Frog?
An interactive benchmark ranks image models by how convincingly they redraw Bufo across ten frog-themed tasks.
Field notes
Each model receives the same Bufo reference image and ten tasks, including prompts in which Bufo can hold or become a coffee. The site averages ten likeness scores into a rating out of 100 and places that beside measured USD-per-image costs. Visitors can make side-by-side picks, but those selections remain local to the browser and do not alter the leaderboard. The page also warns that small score gaps are inconclusive and that its learned scorer can miss the joke.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue