Removing KL Divergence from On-Policy LLM Distillation
September 26, 2026
Research shows that KL divergence is not strictly necessary for effective on-policy distillation of large language models. Simply preserving the update direction toward the teacher for a small subset of highly divergent tokens is sufficient for performance.
HOW THIS AFFECTS YOU
●
builderYou may be able to reduce the complexity and compute required for fine-tuning student models via distillation.
●
researcherThis simplifies the loss function requirements for on-policy distillation tasks.