[arXiv]score: 0.14
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
August 14, 2026
I-SDPO mitigates signal loss in Group Relative Policy Optimization (GRPO) when all sampled responses in a rollout group are incorrect. The method applies instance-level routing to switch between GRPO and privileged self-distillation, using dense token supervision only for unsuccessful groups to prevent biased teacher imitation from hindering reward-improving updates.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy