Test-Time Policy Optimization for Mathematical Reasoning
August 26, 2026
Test-Time Policy Optimization (TTPO) improves reasoning by using an asymmetric objective during inference. It utilizes On-Policy Self-Distillation for rollouts that agree with majority-vote pseudo-labels and employs Grouped RL to penalize disagreeing rollouts, mitigating the corruption caused by incorrect majority votes.
HOW THIS AFFECTS YOU
●
researcherThis offers a way to perform test-time training on reasoning tasks without relying on ground-truth labels.