REINS: Strengthening LLM Refusals via Sparse Autoencoder Steering
August 31, 2026
REINS improves safety by simultaneously suppressing harmful continuation features and enhancing safe refusal features using Sparse Autoencoders (SAEs). This method addresses failures in standard SAE steering where complex prompt wrappers bypass existing inhibitory controls.
HOW THIS AFFECTS YOU
●
researcherYou can utilize this dual-action steering approach to harden models against sophisticated adversarial prompt wrappers.
●
policyThis provides a more robust technical mechanism for implementing safety guardrails at the feature level.