Observed arrival · 2026-09-11
Sane Labs, Teaching Small Models to Read
An independent language-model project training small models from scratch, with knowledge intended to come from supplied context rather than memorized weights.
Field notes
The project exposes unusually detailed training records, including model shape, token counts, hardware, precision, and weight size. Sane-47M and Sane-118M use an in-house BPE tokenizer and 1,024-token context, while the in-training Synth-2 expands to a listed 8k–32k context and uses 12 experts plus a shared expert. The page also separates released work from private or unfinished experiments rather than presenting every project as available.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue