One-Shot On-Policy Distillation Recovers Most Full-Data Gains
September 2, 2026
On-policy distillation (OPD) can achieve most of the performance gains of full-data training using only a single query. The study shows a single query reaches 71.5% state coverage of full-data OPD within the first 100 training steps.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce distillation costs by using highly targeted, single-query training regimes.
●
researcherThis changes how you view the data-minimal limits of student-teacher alignment.