Mechanistic Interpretability of On-Policy Distillation via Swap Readout
September 27, 2026
Using sparse crosscoders and a new 'swap readout' method, this research analyzes how on-policy distillation (OPD) changes a student model's feature usage. The study tracks how internal representations evolve during the transfer of knowledge from a teacher.
HOW THIS AFFECTS YOU
●
researcherThis provides a new tool for understanding the mechanistic changes that occur during post-training distillation.