MedQA-MM reveals reasoning inflation in medical multimodal models
September 4, 2026
Audits across six medical MCQ datasets show that models often rely on text cues, annotations, or device artifacts rather than visual reasoning. In a 13-model panel, full-input accuracy was 62.63%, while text-only and options-only settings achieved 53%, indicating significant reasoning inflation.
HOW THIS AFFECTS YOU
●
researcherYou should account for non-visual cues when evaluating multimodal reasoning capabilities.
●
healthBe cautious of high benchmark scores that may not reflect true visual diagnostic reasoning.