Evaluating Uncertainty Quantification Transfer for LLM Agent Trajectories
August 13, 2026
Evaluation across five LLMs and four tool-use datasets shows that single-turn uncertainty methods transfer unevenly to multi-turn agent trajectories. Token-probability scores depend heavily on temporal aggregation methods, while reflexive self-assessment scores provide the strongest performance in predicting trajectory-level errors.
HOW THIS AFFECTS YOU
●
builderThis suggests caution when using standard token-probability scores to monitor agentic workflows.
●
researcherYou should consider trajectory-level aggregation rather than single-turn metrics when evaluating agent reliability.