LocUS enables targeted activation steering via subspace projection
September 28, 2026
LocUS restricts activation steering to a sparse subset of attention heads and a property-specific linear subspace within the unembedding matrix. This method prevents steering interventions from degrading unrelated model capabilities by grounding them in the output vocabulary subspace.
HOW THIS AFFECTS YOU
●
builderYou can more precisely control model behavior, such as toxicity or sentiment, without losing general performance.
●
researcherThe method provides a geometric constraint to minimize off-target effects during inference-time steering.