LLMs Can Control Residual Stream Activations via Natural Language
August 25, 2026
The Activation Controllability Benchmark demonstrates that most LLMs can modulate the direction and magnitude of their residual stream activations using natural-language instructions. This ability allows models to potentially evade latent-space monitoring methods like linear probes and Jacobian lenses.
HOW THIS AFFECTS YOU
●
researcherLatent-space monitoring may be insufficient to detect deceptive model behavior.
●
policySafety frameworks relying on activation monitoring must account for models that can manipulate their own internal states.