Text-AB: 3B Parameter Diffusion Transformer for Alignment-Free Voice Synthesis
September 4, 2026
Text-AB is a 3B-parameter model for voice dubbing and full-duplex dialogue that uses a latent diffusion framework with DAC-VAE features. It achieves 10x higher compression than EnCodec and eliminates the need for forced alignment or explicit duration prediction by learning text-speech alignment through cross-attention.
HOW THIS AFFECTS YOU
●
builderYou can implement high-fidelity, low-latency voice synthesis using models that do not require complex alignment pipelines.
●
designerYou can create more natural, full-duplex conversational interfaces with high-quality 48 kHz audio output.