Dual-Form ASR Integrates Semantics-Aware ITN for Chinese Speech Recognition
September 4, 2026
Dual-Form ASR (DF-ASR) replaces cascaded ASR and ITN modules with a unified framework trained via LLM-driven spoken-written paired supervision. This method improves the accuracy of written-form transcripts by coupling normalization directly with acoustic-contextual modeling.
HOW THIS AFFECTS YOU
●
builderYou can achieve higher fidelity in transcription products by moving away from decoupled ITN modules.
●
researcherThe LLM-driven generate-and-judge workflow provides a new way to construct supervision data for speech tasks.