[HUGGINGFACE]score: 0.42
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
October 6, 2026
Δ-MOPD improves multi-teacher on-policy distillation by transferring teacher-minus-base logit shifts re-anchored to a student's frozen initialization. This method prevents inherited base preferences from overpowering teacher-specific updates during common-domain or routed-domain distillation, addressing the performance degradation observed when using standard endpoint supervision.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy