Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
September 28, 2026
A training-free pipeline integrates lexical G2P tools with frozen pretrained speech-to-pronunciation models using a left-to-right greedy search. On Japanese corpora, the method reduced Character Error Rate from 0.60--1.40% to 0.04--0.17% compared to text-only baselines.
HOW THIS AFFECTS YOU
●
builderYou can improve TTS training data preparation efficiency without needing costly pronunciation-annotated datasets.
●
researcherThis method demonstrates how text constraints can significantly refine acoustic rescoring during inference.