ReACT-TTS Uses Listener Facial Reactions for Prosody Planning
September 21, 2026
ReACT-TTS is a two-stage framework that uses a one-second pre-response facial sequence to plan utterance emotion and prosody. In dyadic MELD protocol testing, temporal conditioning achieved higher mean macro-F1 and VAD concordance than text-only baselines, with 76% of researchers preferring this method.
HOW THIS AFFECTS YOU
●
researcherYou can leverage temporal visual conditioning to improve prosody modeling in conversational TTS.
●
designerThis enables more emotionally reactive and lifelike AI avatars.