●researcherYou should account for the fact that SFT and RLHF can decouple mechanistic weight edits from observable model behavior.
●policyThis highlights a challenge in ensuring long-term alignment stability as models undergo post-deployment fine-tuning.