SIGNPOST-Bench Evaluates Text-Vision Conflict in MLLMs
August 3, 2026
SIGNPOST-Bench introduces a counterfactual benchmark consisting of 25,555 image variants to test how multimodal models resolve contradictions between visual and textual cues. It uses localized scene-text interventions to measure localization shifts and accuracy when evidence sources disagree.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to evaluate how your MLLM handles conflicting sensory inputs.