FD-VAD enables ASR-free semantic endpoint detection for streaming voice
September 30, 2026
FD-VAD implements an ASR-free streaming endpointer that maps causal audio windows directly to Continue/Stop decisions. The architecture uses a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model to handle turn-taking without transcription dependence.
HOW THIS AFFECTS YOU
●
builderYou can reduce latency and transcription dependencies in full-duplex voice agents by using ASR-free semantic endpointing.
●
designerYou can create more natural, human-like turn-taking interactions by using confidence-gated endpoint commitment.