LLMs Maintain Linear Representations of Contextual Truth in Activation Space
August 5, 2026
Research shows that LLMs encode the truth of context-dependent propositions as linear directions in activation space. These representations persist across different output policies and are susceptible to steering via partner assertions in collaborative tasks.
HOW THIS AFFECTS YOU
●
researcherYou can leverage these linear directions for causal steering experiments regarding truthfulness.