Proposes the Residual Advantage (RA) method to improve post-training reasoning models. RA treats the teacher-student probability residual as a bounded reward, centering it to provide better guidance than standard on-policy distillation.
HOW THIS AFFECTS YOU
●
builderThis technique can help you refine reasoning-heavy models using verifiable reward signals.
●
researcherYou can use RA to improve the efficiency of on-policy distillation in reasoning models.