Mitigating Length-Scaling Tax via Length Self-Distillation
September 29, 2026
Length Self-Distillation (LSD) reduces the length-scaling tax, where RL post-training makes models unnecessarily verbose on solved problems. LSD uses an exponential moving average of the online policy to distill on-policy responses, curbing verbosity without losing accuracy.
HOW THIS AFFECTS YOU
●
builderYou can reduce inference latency and costs by preventing unnecessary response length expansion.
●
researcherThis provides a way to optimize RL objectives for conciseness.