This study demonstrates a method for creating compact, fixed-voice Thai TTS models by using large voice-cloning models as programmable data sources. The approach converts 15-second voice references into high-quality synthetic corpora to train efficient student models for low-resource settings.
HOW THIS AFFECTS YOU
●
builderYou can deploy efficient, small-footprint TTS models for low-resource languages using synthetic teachers.
●
researcherThis method highlights how teacher errors and filtering impact student performance in synthetic training pipelines.