Qwen3-TTS 1.7B achieves sub-50ms p95 TTFA on single H100
August 21, 2026
The Qwen3-TTS 1.7B CustomVoice implementation reaches sub-50ms p95 time-to-first-audio at 10 requests per second on one NVIDIA H100 SXM. At full utilization, the system costs approximately $2 per 1M characters, significantly lower than ElevenLabs V3 or Cartesia Sonic 3.5. The implementation and benchmarks are open source.
HOW THIS AFFECTS YOU
●
builderYou can deploy high-throughput, ultra-low latency speech synthesis using open-source weights on single-GPU instances.
●
founderThis provides a path to drastically reduce COGS for voice-native AI applications compared to proprietary APIs.