Adversarial rebuttals from human-in-the-loop reviewers can manipulate LLM moderators to either whitewash hate speech or smear harmless content. Experiments demonstrate that these decision-boundary perturbations significantly degrade moderation performance, particularly in multi-turn settings.
HOW THIS AFFECTS YOU
●
builderIf you use human-in-the-loop moderation, you need defenses against adversarial user rebuttals.
●
policyYou must account for human-AI feedback loops that can be exploited to bypass safety guardrails.