R2-OPD Filters On-Policy Distillation via Reasoning Progress Awareness
August 18, 2026
R2-OPD addresses mismatches in on-policy distillation where teacher-derived rewards conflict with genuine reasoning progress. The method filters rewards to ensure student models prioritize steps that show actual reasoning advancement rather than mere imitation of teacher outputs.
HOW THIS AFFECTS YOU
●
researcherThis method offers a way to improve distillation quality by decoupling imitation from reasoning capability.