●researcherThis highlights a critical trade-off when using large models to teach smaller ones via distillation.
●policyBe aware that distilling capability into smaller models may inadvertently increase their tendency to output biased content in ambiguous scenarios.