[arXiv]score: 0.18
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
August 12, 2026
SPIEval introduces a human-curated benchmark of 250 tasks designed to evaluate LLMs as mobile assistants managing scattered personal data. The dataset uses 4,335 records across 10 apps and 21 tools to test five cognitive dimensions, including multi-intent decomposition and preference inference. Performance testing across nine models shows significant gaps in handling fragmented information.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy