ContraTalk Benchmark Identifies Text-Biased Reasoning in Audio Models
August 28, 2026
ContraTalk is a new benchmark of 501 questions designed to detect cross-modal disagreement between transcripts and acoustic signals. It forces models to resolve conflicts where lexical content contradicts paralinguistic cues like prosody or emotion to ensure true audio grounding.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate whether your multimodal models are genuinely grounding in audio or just relying on transcripts.