R2-OPD: Reasoning-aware reward filtering for distillation
August 21, 2026
R2-OPD improves on-policy distillation by introducing Reasoning-Progress-Aware Reward Filtering. It filters teacher-derived rewards by comparing them against an independently estimated progress reward to ensure distillation aligns with actual reasoning advancement.
HOW THIS AFFECTS YOU
●
researcherThis addresses the mismatch where teacher feedback may conflict with genuine reasoning progress during policy optimization.