LVLMs Struggle with Interactive Visual Grounding Tasks
August 26, 2026
A new evaluation framework shows that Large Vision-Language Models perform significantly below human baselines when visual grounding requires interactive dialogue to resolve ambiguity. Performance is lowest in proactive question-driven scenarios where models must acquire target information through conversation rather than single-shot descriptions.