Hierarchical Reasoning and Variance-Gated Rewards for Social LLMs
August 7, 2026
The TSR framework decomposes social dialogue into high-level strategic planning and low-level linguistic execution. It utilizes a Linearized Hierarchical Reinforcement Learning algorithm with Variance-Gated Rewards to balance goal completion with strategy adherence on the SOTOPIA benchmark.
HOW THIS AFFECTS YOU
●
researcherThis offers a new method for training agents to handle dynamic social interactions via hierarchical RL.