Budget vs Capability Confounds in LLM Reasoning Traces
September 4, 2026
A study of 178 problem-model cells across MATH problems shows that perceived reasoning breakthroughs are often artifacts of token budget rather than inherent model capability. Using restart-controlled truncation probes, researchers found that continuing a model's own prefix beats restarting in 9 of 9 matched-budget cases, suggesting compute-heavy trajectories are often just budget management.
HOW THIS AFFECTS YOU
●
researcherYou should use restart-controlled controls to distinguish between a model's prefix value and simple continuation budget.