TaxoBench reveals deep research agents retrieve only 21% of essential papers
August 7, 2026
TaxoBench evaluates the ability of research agents to retrieve and organize papers into expert taxonomies. Testing 7 agents shows the top performer retrieves only 20.92% of expert-cited papers, and no standard configuration matches the average expert taxonomy depth of 4.86.
HOW THIS AFFECTS YOU
●
builderCurrent agents struggle with comprehensive literature retrieval and structural organization tasks.
●
researcherYou should account for a high synthesis gap when evaluating agentic research workflows.