Interrogating Multimodal Diffusion Transformers via LLM Bottlenecks
October 6, 2026
A lightweight bottleneck network can map intermediate multimodal attention tokens from Diffusion Transformers into a frozen LLM's input space. This allows an LLM to answer natural language questions about emerging image semantics and fine-grained details directly from hidden representations.
HOW THIS AFFECTS YOU
●
researcherYou can now interpret the hidden semantic space of multimodal diffusion models.
●
designerYou can gain better control over image generation by understanding how text tokens evolve during denoising.