Multi-Agent Debate Fails to Outperform Self-Consistency at Scale
September 30, 2026
An analysis of 23 small language models shows that multi-agent debate (MAD) does not derive gains from cognitive diversity in personas, temperature, or model identity. At matched budgets, MAD ties or loses to self-consistency sampling despite 1.6x higher wall-clock and 3.4x higher token costs.
HOW THIS AFFECTS YOU
●
builderYou should consider self-consistency over multi-agent debate to optimize for token cost and latency.
●
researcherThe findings challenge the assumption that agentic diversity is a primary driver of reasoning gains.