DASH Improves Reasoning Models via Divergence-Adaptive Self-Distillation
August 7, 2026
DASH enhances on-policy self-distillation (OPSD) for reasoning models by accounting for the temporal structure of teacher-student divergence. Rather than using static coefficients, it adjusts supervision based on the divergence history and position within the autoregressive rollout to better exploit temporal context.
HOW THIS AFFECTS YOU
●
researcherYou can improve the efficiency of RLVR by using divergence-adaptive horizons instead of uniform token-level supervision.