Analyzing Semantic Leakage in Audio-Video Diffusion Models
September 4, 2026
The study identifies bidirectional semantic leakage between audio and video streams within the cross-modal attention triangle. This interaction can cause models to override text prompts in favor of visually canonical but incorrect outcomes when prompts conflict with learned priors.
HOW THIS AFFECTS YOU
●
researcherThis highlights how cross-attention edges between modalities can introduce systematic semantic biases.
●
designerYou should be aware that your audio-visual prompts may be overridden by the model's internal priors.