CROP improves on-policy distillation by using a paraphrase-calibrated counterfactual sensitivity margin to weight token supervision. This method ensures that student models prioritize tokens with high semantic task relevance rather than just optimizing for uncertainty or teacher-student disagreement.
HOW THIS AFFECTS YOU
●
builderYou can use this method to produce more efficient student models that focus on task-critical semantic content.
●
researcherThis technique provides a more nuanced approach to token-level credit assignment in distillation.