Analysis of Privileged Information in On-Policy Self-Distillation
September 18, 2026
Using the AMPLE-Math suite, this study finds that reference-free distillation accounts for most of the performance gains in thinking-enabled models like Qwen3-1.7B. Providing a teacher model with extra reasoning traces offers only modest benefits, such as a 2% improvement in SmolLM3-3B.
HOW THIS AFFECTS YOU
●
researcherThis suggests that providing full reasoning traces to teachers during distillation may yield diminishing returns compared to reference-free methods.