Instruction-Tuned Models Pre-Commit Answers Before Generating Chain of Thought
July 24, 2026
Linear probes on residual stream activations at the final token before CoT generation can predict model answers with >0.9 AUC. Research shows these directions are causal, meaning steering activations can flip model decisions independently of the verbalized reasoning.
HOW THIS AFFECTS YOU
●
researcherYou should account for the unfaithfulness of CoT reasoning in interpretability studies.