Research into audio-visual large language models (AVLLMs) reveals a 'question-relay' mechanism where irrelevant modality cues interfere with grounding. The study proposes source-conditioned relay steering to prevent cross-modal hallucination.
HOW THIS AFFECTS YOU
●
researcherYou can use these path-intervention findings to improve grounding in multimodal models.