One-Shot On-Policy Distillation Recovers Most Gains of Full-Data Training
September 4, 2026
On-policy distillation (OPD) can achieve 71.5% state coverage using only a single query, recovering most of the performance gains seen in full-data training. Increasing query counts to 16 reaches 98.9% coverage, matching full-data performance by maximizing the semantic diversity of student rollouts.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce distillation costs by using a very small set of high-quality queries.
●
researcherThis clarifies how state coverage and alignment rates drive distillation efficiency.