EmoRSS Uses Activation Steering to Reduce LLM Over-Refusal
October 6, 2026
EmoRSS mitigates the tendency of LLMs to refuse benign requests triggered by emotional language. The method uses sparse autoencoder features to steer the refusal subspace, preserving safety responses for harmful queries while increasing receptivity to non-harmful emotional input.
HOW THIS AFFECTS YOU
●
researcherYou can apply activation-steering via SAE features to refine safety alignment.
●
policyThis provides a technical pathway to reduce the friction caused by over-sensitive safety guardrails.