MME-Safety provides a four-dimensional annotation schema to evaluate MLLM vulnerabilities across risk scenarios, harm severity, and modality-specific stealth levels. Zero-shot evaluations of 17 state-of-the-art models reveal how cross-modal inputs bypass traditional unimodal safety filters.
HOW THIS AFFECTS YOU
●
researcherYou can use this hierarchical framework to identify specific structural weaknesses in multimodal defensive behaviors.
●
policyThis provides more granular metrics for assessing the real-world risks of multimodal AI deployments.