Initialization Anchoring Weakness in Feedback-Based Agent Planning
September 25, 2026
Feedback-based agents exhibit an initialization anchoring weakness where the first feedback round corrects 46% of adversarial directions, but effectiveness drops to 7% by the third round. The proposed InitAnchor framework exploits this via attacker-controlled materials that leverage contextually plausible shifts and persistent trajectory directions.
HOW THIS AFFECTS YOU
●
researcherYou should account for decaying corrective efficacy in multi-turn agentic feedback loops.
●
policyThis reveals a specific vulnerability in agentic planning that could be exploited through external context injection.