Tokenization Vulnerabilities in Knowledge Editing and Unlearning
September 25, 2026
Alternative valid tokenizations can bypass localized model editing and unlearning by inducing different computational trajectories for the same input string. This finding demonstrates that current security evaluations of machine unlearning are insufficient when tokenization is treated as a benign preprocessing step.
HOW THIS AFFECTS YOU
●
researcherYou must account for tokenization variance when testing the robustness of knowledge-editing techniques.
●
policyThis highlights a critical security gap in how we guarantee the removal of sensitive data from open-weight models.