The DRIFT method identifies vulnerabilities in flow-matching vision-language-action models by applying adversarial patches to the robot gripper. The research demonstrates that attacking only the first denoising step is more effective and efficient at derailing robot trajectories than attacking the entire window.
HOW THIS AFFECTS YOU
●
researcherYou should account for first-step denoising vulnerabilities when training robust VLAs.
●
policyThis highlights security risks in autonomous robotic systems using flow-matching.