Semantic-shift attacks bypass safety filters by replacing harmful terms with benign alternatives. This method optimizes the surrounding context to maximize the model's likelihood of reinterpreting these benign terms as their harmful counterparts, enhancing jailbreak success rates.
HOW THIS AFFECTS YOU
●
researcherYou can use context optimization to better evaluate the robustness of safety alignment.
●
policyThis highlights how subtle linguistic shifts can circumvent existing guardrails.