[HUGGINGFACE]score: 0.42
Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
October 4, 2026
Seven 7B-8B parameter agents correctly identify failing tool outputs as useless in 97-100% of cases but rarely use these judgments to terminate execution. While prompt cues influence stop timing, they do not improve the accuracy of the decision to quit. Incorporating explicit stopping rules or call costs is necessary to prevent redundant queries.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy