Observed arrival · 2026-09-15
NLASmith Turns AI Interpretability Into a Repeatable Experiment
An evaluation framework for running and comparing Natural Language Activation experiments across datasets, token positions, evaluators, and metrics.
Field notes
The framework fixes the experimental unit before execution: a prompt dataset, a token-position rule, a Neuronpedia activation lookup, and an evaluator schema. It distinguishes per-example artifacts from aggregate measures such as presence rate, mean score, and category mix, allowing runs on the same prompt set to be compared. The site also explicitly limits its claim: NLASmith is presented as infrastructure for testing hypotheses, not as a new interpretability method or a newly trained autoencoder.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue