CRAG-MM-Diagnostics isolates VQA failures in grounding and retrieval
July 24, 2026
CRAG-MM-Diagnostics introduces a diagnostic benchmark for Knowledge-Intensive Visual Question Answering that decomposes failures into language-based grounding, object identification, and knowledge retrieval. The framework uses metadata like target ROIs and visual complexity scores to analyze both parametric and retrieval-augmented vision-language models.
HOW THIS AFFECTS YOU
●
researcherYou can use this to pinpoint whether VLM errors stem from visual perception or knowledge retrieval gaps.