On-Policy Power Distillation Improves LLM Reasoning
October 4, 2026
On-policy power distillation (OPPD) trains models to generate answers following a sharpened probability distribution. This allows models to capture the reasoning benefits of sequence-level power distribution in a single generation pass without extra sampling.
HOW THIS AFFECTS YOU
●
researcherThis offers a more efficient training alternative to expensive sequential Monte Carlo sampling for reasoning tasks.