Task Difficulty Governs Causal Load-Bearingness of Chain-of-Thought
September 23, 2026
Continuation-based causal testing reveals that models bypass reasoning on easy tasks but propagate errors on difficult ones. Testing across Gemma-2-9B, Llama-3.1-8B, and DeepSeek-R1 shows that error propagation increases 16x when moving from GSM8K to BBH multitask arithmetic.
HOW THIS AFFECTS YOU
●
researcherYou can use this testing framework to distinguish between mechanistic faithfulness and mere behavioral imitation in CoT reasoning.