GUARD Method Implements Natural Forgetting for Reasoning Models
September 21, 2026
GUARD addresses unsafe rationales in Large Reasoning Models (LRMs) by distilling a coherent, non-disclosing Chain-of-Thought (CoT) followed by a stable refusal. This prevents the hallucinated substitutes or malformed boundaries common in standard machine unlearning objectives.
HOW THIS AFFECTS YOU
●
researcherYou can use guided distillation to create more stable unlearning trajectories in CoT-based models.
●
policyThis provides a more robust method for removing sensitive reasoning traces from large models.