A new method improves speech language model performance on dialects by synthesizing pseudo-dialect speech using standard-language TTS and LLM-generated text. This approach requires zero real dialect speech and uses intermediate standard-text prediction to bridge semantic gaps.
HOW THIS AFFECTS YOU
●
builderYou can improve the robustness of your speech models for low-resource dialects without collecting new audio data.
●
researcherThis provides a zero-shot data augmentation technique for dialect-robustness in SLMs.