FACA Improves Multi-Turn Agent Performance via Feedback-Aware Credit Assignment
August 19, 2026
FACA aligns agent reactions with specific dialogue segments by deriving locally normalized reaction advantages. In nine-domain simulations, this method improves the average tau-family metric by up to 10.22 percentage points compared to standard outcome-only Interactive GRPO.
HOW THIS AFFECTS YOU
●
builderThis provides a way to reduce errors in interactive tool-use agents during long-running user sessions.
●
researcherYou can better optimize multi-turn agents by using local feedback signals instead of relying solely on terminal rewards.