MMOOC Benchmark for Multimodal Large Language Model Robustness
July 31, 2026
MMOOC is a 41K image-question pair benchmark designed to evaluate MLLM refusal and answering abilities under context shifts. It distinguishes between unanswerable out-of-context (OOC) questions and answerable shifted in-context (Shifted IC) questions.
HOW THIS AFFECTS YOU
●
researcherYou can use this to better evaluate how MLLMs handle subject-level vs. context-level shifts.
●
policyThis benchmark helps identify reliability gaps in how models refuse or accept out-of-context queries.