Refined SFT Token Reweighting for Feature Reversal and Extrapolation
September 26, 2026
This research identifies that standard SFT token reweighting cannot reverse learned harmful features and proposes a new method using a stable SFT delta as a reference frame. This allows for more controlled feature extrapolation and reversal during the fine-tuning process.
HOW THIS AFFECTS YOU
●
builderYou can potentially use these techniques to unlearn specific behaviors more effectively.
●
researcherThis provides a more mathematically grounded view of the optimization trajectory in SFT.