MC-CXR Benchmark Evaluates VLM Robustness to Conflicting Clinical Context
August 26, 2026
The MC-CXR benchmark comprises 2,522 instances designed to measure how vision-language models respond when clinical context, such as prior reports, conflicts with chest X-ray images. It introduces metrics like switch-to-wrong rate to identify context-induced disruption in medical VLMs.
HOW THIS AFFECTS YOU
●
researcherYou can use this to evaluate if your VLM's vision-only decisions are being corrupted by text context.
●
healthThis highlights critical reliability risks in deploying VLMs for clinical diagnostic support.