LSPD: Sample-Efficient LLM Reasoning via Least-Square Policy Distillation
September 27, 2026
LSPD connects on-policy distillation to KL-regularized policy optimization to introduce optimistic exploration and off-policy data reuse. The framework achieves an O(log K) regret bound and improves rollout efficiency by learning from previously collected trajectories.
HOW THIS AFFECTS YOU
●
researcherYou can improve distillation efficiency by treating the objective through an RL lens.