Mechanistic Analysis of Multilingual Refusal Circuits in MoE
August 11, 2026
Investigation of the sarvam Indic-multilingual MoE model shows that safety refusal is a late-stage process where a specific mixture-of-experts writer is controlled by an attention opposer. While harm detection is language-invariant in mid-network layers, the refusal writing occurs late in the generation process.
HOW THIS AFFECTS YOU
●
researcherYou can target the attention opposer circuit to more efficiently control multilingual refusal behaviors.
●
policyThis explains why safety alignment is often unevenly applied across high-resource and low-resource languages.