MMDiff Framework for Multimodal Feature Discovery and Control
August 9, 2026
MMDiff uses multimodal sparse autoencoders (SAEs) to identify, isolate, and control internal features in MLLMs. By diffing base-LM SAEs against multimodal-adapted versions, the framework allows for precise discovery of features altered during multimodal training.
HOW THIS AFFECTS YOU
●
builderYou can gain more granular control over multimodal model behaviors via feature-level interfaces.
●
researcherYou can use this to audit how multimodal training changes the internal representations of an LLM.