V-DEAL Framework Identifies Safety De-Calibration in Video LLMs
July 24, 2026
The V-DEAL diagnostic framework reveals that Video LLMs exhibit an understanding-refusal coupling failure, where benign queries paired with harmful videos achieve higher attack success rates than explicit queries. Tests on six models showed that despite 81% accuracy in recognizing harmful content, average attack success rates remained at 48.33%.
HOW THIS AFFECTS YOU
●
researcherYou should investigate the decoupling of perception and refusal in multimodal safety alignment.
●
policyThis highlights a significant vulnerability in how video models handle safety constraints during deployment.