Causal Tracing Identifies Safety Judgment Locations in Vision-Language Models
October 7, 2026
Using the SSU-Bench dataset, researchers mapped where vision-language models integrate image and text to form safety verdicts. Findings show that interventions at changed input positions influence earlier decoder layers, while final-token interventions impact later layers.
HOW THIS AFFECTS YOU
●
researcherThis provides a method for tracing how multimodal safety judgments are synthesized within model layers.
●
policyUnderstanding the internal mechanics of safety decisions is critical for auditing multimodal model alignment.