LastOPD Mitigates Collapse in Latent On-Policy Distillation
September 22, 2026
LastOPD investigates why latent supervision in on-policy distillation causes performance collapse in Qwen3 models. The research identifies that while latent alignment improves, it can lead to degraded behavior and accuracy drops after initial gains.
HOW THIS AFFECTS YOU
●
researcherThis serves as a warning that maximizing latent alignment with a teacher model does not guarantee improved functional performance.