Agent Retrieval Bench evaluates file-level context retrieval for coding agents
July 26, 2026
Agent Retrieval Bench moves beyond patch-level evaluation to test the upstream capability of coding agents to find relevant repository files. The benchmark includes 427 samples across five tasks like code2test and trace2code, prioritizing workflow-based relevance over simple semantic similarity.
HOW THIS AFFECTS YOU
●
builderYou can specifically measure and improve the retrieval component of your coding agent's architecture.
●
researcherYou can evaluate agents based on real-world repository navigation rather than just final code output.