●builderYou can use this technique to distill smaller, faster models that maintain high reasoning reliability without reward hacking.
●researcherThis addresses a fundamental flaw in KL-based distillation where local logical errors are masked by global teacher similarity.