Researchers identified trait-induced safety variation, where system prompt traits perturb safety representations in a low-dimensional subspace. They propose Trait-Invariant Safety Tuning to ensure models maintain consistent refusal and compliance behaviors regardless of the persona or traits assigned in the system prompt.
HOW THIS AFFECTS YOU
●
researcherYou can use these refusal-based metrics to measure how system prompts destabilize safety alignments.
●
policyThis work highlights how easily safety guardrails can be bypassed or altered via persona manipulation.