The BATON framework improves LLM agent performance by combining Bayesian Feedback Attribution to attribute environment feedback within trajectories and Trajectory Mass Normalization to equalize optimization weight across batches. Testing on ALFWorld, WebShop, and SearchQA shows consistent gains across various model scales.
HOW THIS AFFECTS YOU
●
researcherYou can apply this dual-axis approach to optimize agentic reinforcement learning workflows.