DIAL-OPD Method for Efficient On-Policy Distillation
October 9, 2026
DIAL-OPD improves on-policy distillation by weighting rewards with the logarithmic mean of teacher and student probabilities to select high-value tokens. Testing on seven mathematical reasoning benchmarks shows that training on fewer, strategically selected tokens can outperform full-vocabulary supervision.
HOW THIS AFFECTS YOU
●
builderThis offers a path to reduce training costs while potentially improving reasoning performance.
●
researcherYou can optimize distillation training by moving away from full-token supervision toward value-based token selection.