ENDOPROMPT method degrades instruction-tuned model utility by 26.8 percent
September 25, 2026
ENDOPROMPT is a white-box attack that generates utility-degrading prefixes to reduce task performance without triggering safety filters. Tests across four models showed a mean utility drop of 26.8 percentage points through local search and preference fitting.
HOW THIS AFFECTS YOU
●
researcherYou should account for non-harmful utility degradation when evaluating instruction-tuned model robustness.
●
policyYou must consider attacks that bypass safety alignment by simply neutralizing model utility.