Box2-Bench Evaluates LLM Robustness to Unreliable Agent Workflows
October 1, 2026
The Box2-Bench framework measures how models regulate reliance on external guidance by varying workflow reliability. Results show frontier models remain vulnerable to misleading guidance, though counterfactual supervised fine-tuning and outcome-based RL can improve robustness.
HOW THIS AFFECTS YOU
●
builderYou should implement safeguards when using agentic workflows, as even frontier models struggle to override bad guidance.
●
researcherThis provides a method to isolate a model's ability to selectively use external tools and workflows.