Semi-OPD Outperforms On-Policy Distillation in Most Cases
October 9, 2026
Semi-OPD, which distills from offline student rollouts, outperforms traditional on-policy distillation (OPD) in 14 of 17 tested teacher-student pairs. This method achieves up to 13.6% higher accuracy and an 11.4x speedup in training efficiency.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce distillation costs and training time using offline student rollouts.
●
researcherThe effectiveness of on-policy sampling is highly dependent on the token overlap ratio between models.