SAKI Method Optimizes On-Policy Distillation via Maximal-Coupling-Routed Supervision
September 30, 2026
SAKI improves on-policy distillation by routing token-level supervision based on whether a student accepts or corrects a teacher's rollout. It uses maximal coupling and KL-constrained interpolation to ensure that correction positions receive direct supervision while maintaining the exact-q trajectory distribution.
HOW THIS AFFECTS YOU
●
researcherYou can use this method to reduce state mismatch during student training by more intelligently allocating teacher supervision.