[arXiv]score: 0.40SearchAuditBench and SearchAuditor for Debugging Long-Horizon AgentsAugust 7, 2026SearchAuditBench provides 1,243 expert-annotated failed trajectories, averaging 73.1 messages and 65.1K tokens, to evaluate LLM auditing capabilities. The SearchAuditor framework localizes, attributes, and repairs reasoning errors within complex, multi-step web search trajectories.HOW THIS AFFECTS YOU●builderYou can use this benchmark to evaluate how effectively your agents can self-correct during long-horizon tasks.●researcherThis provides a large-scale, annotated dataset for studying error propagation in reasoning agents.read original ↗arxiv.orgDAILY DIGEST_all newsbuilderresearcherfounderinvestordesignerpolicyhealthsubscribe →you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy← back to feed