BM25 Outperforms Agentic RAG at Scale in Corpus Expansion Study
July 30, 2026
A study of 28 corpus tiers shows BM25 defines the low-cost Pareto frontier and leads accuracy from mid-scale onward. While File-System Agents match performance at small scales, they consume 39 times more query tokens than BM25.
HOW THIS AFFECTS YOU
●
builderYou can optimize for cost and accuracy by using BM25 as your retrieval backbone for large-scale document sets.
●
founderThis suggests that complex agentic RAG may not be a competitive advantage for high-volume, large-scale retrieval tasks.