Deriving steering vectors from concept tokens via Jacobian space inversion
September 6, 2026
Inverting the Jacobian lens allows for the extraction of general activation steering vectors using only concept-related tokens. Testing on Qwen3-1.7B shows success for simple behaviors like capitalization changes, though the method remains brittle for complex linguistic patterns and increases hallucination risks.
HOW THIS AFFECTS YOU
●
builderYou can use this for low-cost behavioral steering, but expect instability in complex reasoning tasks.
●
researcherYou can explore J space to automate the discovery of steering vectors from specific token sequences.