[HUGGINGFACE]score: 0.47
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
October 5, 2026
Cross-tokenizer on-policy distillation succeeds by focusing on strict 1:1 alignment, as shared vocabularies retain most probability mass despite mismatch. Restricting reverse KL divergence to a student-selected top-16 subset of the shared vocabulary maintains performance across mathematical reasoning and code generation tasks.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy