MissionBench: Zero-Shot Aerial MLLM Evaluation for 3D Tasks
July 27, 2026
MissionBench evaluates Multimodal Large Language Models on 120 long-horizon missions in simulated 3D aerial environments. Benchmarks show current top-performing models achieve less than 35% success rate compared to 84.4% for humans, despite scaling gains in zero-shot embodied reasoning.
HOW THIS AFFECTS YOU
●
builderUse this to benchmark how well general-purpose MLLMs handle complex, multi-step robotic planning.
●
researcherThis provides a new standardized metric for evaluating embodied agent reasoning.