Isolating Category-Specific Harm Signals in Instruction-Tuned LLMs
September 18, 2026
Researchers isolated category-specific harm residuals by removing shared general harmfulness representations across 11 risk categories in three instruction-tuned models. Activation steering using these orthogonal residuals shows that the ability to induce refusal varies significantly depending on the specific harm category.
HOW THIS AFFECTS YOU
●
researcherYou can use these category residuals for more granular activation steering to target specific safety risks.
●
policySafety evaluations should move beyond general harm detection toward category-specific mitigation strategies.