Introduces AnTrap, a benchmark for testing Android GUI agents against runtime anomalies like pop-ups and action misuse. Testing 16 models reveals significant performance degradation across the board, even for top-tier agents.
HOW THIS AFFECTS YOU
●
builderYou need to account for these runtime perturbations when deploying agents in real-world Android environments.
●
researcherThis provides a structured taxonomy and training pipeline to improve agent reasoning in adversarial settings.