Weak-to-Strong On-Policy Distillation for Improving Stronger Students
July 27, 2026
W2S-OPD allows a strong student model to improve by distilling from multiple smaller, weaker models. It uses a proxy teacher constructed in logit space from a contrast pair of positive and negative models to align the student's token-level distribution.
HOW THIS AFFECTS YOU
●
builderThis provides a potential path for improving model performance without relying on frontier-class teachers.
●
researcherYou can now explore distillation techniques that don't require a larger teacher model.