Observed arrival · 2026-09-02
LocoMusa: A Benchmark for Thinking With Local LLMs
A curated benchmark that compares local language models on ideation, reframing, steelmanning, and multi-turn dialogue.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
LocoMusa evaluates local models through blind pairwise comparisons rather than a single intelligence score. Its listed task areas cover ideation, reframing, steelmanning, and multi-turn dialogue, while the published ratings use a Bradley–Terry model normalised against frontier anchors averaging 100. The linked repository contains the tasks, evaluation harness, judge, and methodology; the homepage also points to a JSON leaderboard.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue