R^3-Bench evaluates how LLMs allocate limited computation across shared budgets in math, programming, and reasoning tasks. The benchmark shows that models struggle to maximize expected value when forced to balance multiple problems under a single constraint.
HOW THIS AFFECTS YOU
●
researcherYou can use this to study how models perform in constrained, multi-task environments rather than per-task independent budgets.