Limitations of Outcome-Only Evaluation for Agent Trajectories
September 2, 2026
Outcome-only LLM judges capture only 45% of silent faults—where the final answer is correct despite errors in the process—while incorrectly flagging 33% of successful trajectories. Step-rubric judges are required to achieve higher silent fault recall.
HOW THIS AFFECTS YOU
●
builderYou should avoid relying solely on final-answer correctness for evaluating agent reliability in production.
●
researcherEvaluation frameworks must incorporate step-by-step rubrics to detect subtle process failures.