Per-Call-Site Evaluation Shows Small Models Outperform Frontier Models in Specific Tasks
October 8, 2026
An evaluation of a deployed home-automation system shows that model capability is not uniform across different task types. In a study of 2,520 calls, some 4B models performed worse than 2B models at grounded actuation, suggesting that routing specific tasks to smaller, specialized models is more efficient than using a single frontier model.
HOW THIS AFFECTS YOU
●
builderYou can optimize latency and cost by routing specific intent or planning calls to smaller, task-optimized models.
●
founderThis demonstrates that you don't always need expensive frontier APIs to build reliable agentic products.