NVAlign Optimizes Non-Verbal Control in Flow-Matching TTS
September 24, 2026
NVAlign is a direct-gradient post-training framework designed to improve non-verbal vocalization (NVV) tag following in continuous autoregressive flow-matching text-to-speech models. It uses a frozen NV-ASR reward model to enable efficient gradient backpropagation through the flow-matching sampler.
HOW THIS AFFECTS YOU
●
builderYou can implement more precise control over emotional and non-verbal cues in speech synthesis.
●
designerThis allows for more expressive and nuanced auditory user experiences in generative media.