Llama 3.1 8B Encodes Partisan Identity as Geometric Directions
September 9, 2026
Mechanistic analysis reveals that partisan identity is stored as a locatable geometric direction within Llama 3.1 8B, which alignment training masks rather than deletes. The study demonstrates how model training cutoffs can be used to steer these latent political biases.
HOW THIS AFFECTS YOU
●
researcherYou can use steering experiments to probe and manipulate latent political alignments in LLMs.
●
policyYou should consider how invisible framing and encoded biases affect democratic information access.