Simple-OPD: Improving On-policy Distillation via LoRA Warm-up
August 10, 2026
Simple-OPD enhances on-policy distillation by using LoRA-based warm-up on teacher-generated chain-of-thought data before full distillation. Findings show that warm-up transfers teacher-compatible thinking patterns and that near-saturation LoRA training balances in-domain adaptation with OOD generalization better than full-parameter SFT.
HOW THIS AFFECTS YOU
●
builderYou can use this plug-and-play initialization to improve student model performance during distillation.
●
researcherThis clarifies the role of warm-up in transferring cognitive patterns rather than just factual correctness.