Audio-Visual LLMs Fail Under Cross-Modal Conflict Due to Prior Dominance
August 31, 2026
Analysis of VideoLLaMA 2-7B-AV and InternVideo2 reveals models struggle when audio and video inputs conflict, often reverting to internal output priors. InternVideo2 saw a 32.3% accuracy drop and 17.3% instruction-following failure under conflict, with mechanistic analysis locating decision commitment around layer 25.5.
HOW THIS AFFECTS YOU
●
researcherYou should account for late-layer prior dominance when evaluating multimodal reasoning capabilities.