AgentJudgeBench Evaluates LLM Judge Reliability in Tool-Calling Workflows
August 28, 2026
AgentJudgeBench tests LLM-as-a-judge reliability on 3,808 tool-calling instances across six DAG topologies. Findings show judge alignment degrades significantly with task difficulty and converges to a 77-82% accuracy ceiling on hard queries without ground truth.
HOW THIS AFFECTS YOU
●
researcherYou should account for a structural performance ceiling when using LLMs to evaluate complex agentic workflows.