OPSRD Enables On-Policy Self-Distillation Using Expert Role Prompting
October 1, 2026
On-Policy Self-Role Distillation (OPSRD) uses a frozen, role-prompted teacher to provide conditional distributions for a role-free student. The method applies teacher-weighted forward KL targets to expose and teach the student useful next-token preferences that it would not otherwise sample.
HOW THIS AFFECTS YOU
●
builderThis offers a lightweight way to improve model reasoning by transferring role-based preferences during training.
●
researcherYou can use this technique to distill specialized expert behaviors into smaller models without reference solutions.