Forget-Retain Alignment Gap Predicts LLM Relearning Robustness
August 27, 2026
The Forget-Retain Alignment Gap (FRAG) identifies unlearning robustness by measuring how updates align with forget-critical versus retain-critical weights. This training-free predictor outperforms global weight-space distance metrics in identifying updates vulnerable to knowledge revival via brief fine-tuning.
HOW THIS AFFECTS YOU
●
researcherYou can use FRAG to assess if unlearning updates are selective enough to prevent rapid knowledge re-emergence.