Observed arrival · 2026-09-02
jBlaze, Behavioral Surgery for Language Models
A representation-engineering toolkit that modifies targeted behaviors in open-weight language models without fine-tuning or training data.
Field notes
jBlaze describes a weight-space approach in which behavioral directions can be combined into purpose-built variants instead of being learned through additional datasets or fine-tuning. The page lists more than 25 directions, including refusal, hedging, verbosity, sycophancy, truthfulness, and adversarial resistance, and says its catalog spans 55+ models across six families. Several future directions are explicitly labeled in development, while implementation details and benchmarks are not shown in the extracted page.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue