EraseSAE: Sparse Autoencoder-Based Concept Erasure for Text-to-Video Diffusion
September 4, 2026
EraseSAE targets fine-grained concept removal in DiT-based text-to-video models by intervening at the monosemantic feature level. The framework uses sparse autoencoders to decompose attributes, enabling surgical erasure of specific semantics without degrading overall model generation quality.
HOW THIS AFFECTS YOU
●
researcherYou can use sparse autoencoders to target specific monosemantic features rather than coarse weights.
●
policyThis provides a more precise technical mechanism for enforcing safety and copyright compliance in generative models.