An analysis of failure-based knowledge distillation shows that injecting rule atoms derived from a single model's errors can be unstable. These rules can act as model-specific reasoning patches that confuse other models or cause incorrect answer flips.
HOW THIS AFFECTS YOU
●
researcherYou should be cautious when using failure-derived rules for distillation, as they may lack generalizability.