Reasoning Models Fail to Optimize Shared Token Budgets Across Multiple Questions
August 11, 2026
Frontier reasoning models fail to distribute limited test-time compute strategically when faced with multiple tasks. Evaluation shows models behave as greedy sequential solvers, prioritizing questions by presentation order and front-loading effort rather than maximizing scores based on task difficulty or point values.
HOW THIS AFFECTS YOU
●
builderExpect suboptimal performance in multi-turn or multi-task workflows where total inference cost is constrained.
●
researcherYou should account for budget allocation inefficiencies when evaluating reasoning capabilities.