Contrastive Projection enables clearer reading of transformer internal states
September 10, 2026
By differencing hidden states of closely matched prompts, Contrastive Projection cancels generic components to surface specific model distinctions. This training-free method successfully traces specific MLP-to-attention chains and recovers steering vectors without needing activation patching.
HOW THIS AFFECTS YOU
●
researcherYou can more accurately interpret model internals by analyzing the delta between prompts rather than individual hidden states.