Negative Self-Distillation (NSD) prevents LLMs from suppressing uncertainty by optimizing them to avoid flawed reasoning rather than imitating privileged solutions. This approach addresses the performance degradation caused by traditional On-Policy Self-Distillation during complex reasoning tasks.
HOW THIS AFFECTS YOU
●
researcherThis changes how you approach self-improvement training for reasoning-heavy models.