FailBench: VLM Reliability in Robot Task Evaluation
September 4, 2026
A benchmark containing 2,197 manipulation attempts across 14 sources to test Vision-Language Models on robot failure detection. Results show the best model achieves only 0.77 mean balanced accuracy, with performance dropping below 0.60 on contact-intensive assembly tasks.
HOW THIS AFFECTS YOU
●
builderYou should account for low reliability when using VLMs to judge physical robot outcomes.
●
researcherYou can use this to evaluate VLM generalization in robotic failure modes.