LifePlanner Benchmark for Geo-spatial Agent Reasoning
August 27, 2026
LifePlanner evaluates LLM agents on geo-spatial planning tasks by integrating large-scale social media data via an MCP toolset. Results show frontier models struggle with complex planning, with pass rates dropping to 40.2% due to failures in evidence acquisition and constraint integration.
HOW THIS AFFECTS YOU
●
builderYou should focus on improving tool use and multi-constraint reasoning for location-aware agent applications.
●
researcherYou can use this benchmark to evaluate how agents handle noisy, real-world social signals in planning.