Agent orchestration scaling returns vary by task and architecture
September 30, 2026
An evaluation of eight orchestration architectures across 7-9B models shows that scaling from 3 to 30 agent calls yields up to 17 points on arithmetic benchmarks but minimal gains on MMLU or GPQA. Proposer-Critic scales most effectively for arithmetic tasks but lacks generalist superiority.
HOW THIS AFFECTS YOU
●
builderYou should select orchestration architectures based on specific task types rather than general accuracy.
●
founderThis suggests that increasing agent count is not a universal solution for all reasoning capabilities.