Skip to the card

Card 472 of 9992026-09-02 issue

Observed arrival · 2026-09-02

LocoMusa: A Benchmark for Thinking With Local LLMs

locomusa.org Observed source
Editorial interest 78/100 Selection signal · not a rating of the site

A curated benchmark that compares local language models on ideation, reframing, steelmanning, and multi-turn dialogue.

Landing page captured for the 2026-09-02 issue.

Field notes

LocoMusa evaluates local models through blind pairwise comparisons rather than a single intelligence score. Its listed task areas cover ideation, reframing, steelmanning, and multi-turn dialogue, while the published ratings use a Bradley–Terry model normalised against frontier anchors averaging 100. The linked repository contains the tasks, evaluation harness, judge, and methodology; the homepage also points to a JSON leaderboard.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

A Commons Made of Smoke

312,505 arrived 1,000 judged 999 catalogued Enter the complete issue
locomusa.org

Landing page observed 2026-09-02. The live site may have changed.