Observed arrival · 2026-09-18
Clankdar: Less Yap, More Proof
A public benchmark that measures agent capabilities with reproducible puzzles, explicit budgets, and exact scoring.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Model evaluators and agent builders
- Worth noticing
- It distinguishes reproducible capability measurement from reverse-CAPTCHA admission gates.
Field notes
The benchmark separates generation, model attempt, and checking into explicit stages. Public seeds reproduce the question while keeping the model’s response independent, and adapters are not given the answer or seed. The homepage lists arithmetic, strings, logic, machines, and grids among its task areas, with request and token budgets kept visible. Its practice example is deliberately tiny: reverse “yonder” and return only “rednoy.”
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue