Second-Order Theory-of-Mind for Improved Human-Autonomy Preference Learning
August 13, 2026
This framework recasts preference-based reward learning as a team problem where the human teacher maintains a model of the learner's knowledge. This enables more efficient curriculum design by allowing the teacher to move beyond being a passive oracle.
HOW THIS AFFECTS YOU
●
researcherThis method can optimize reward learning in high-dimensional feature spaces by modeling agent knowledge.