YODAS v3 Releases 1.1 Million Hours of Multilingual Stereo Speech
September 25, 2026
YODAS v3 provides a weakly-labeled, 48kHz stereophonic speech corpus covering 147 languages. It features 22 languages with over 10,000 hours of data, serving as the largest open high-fidelity multilingual dataset.
HOW THIS AFFECTS YOU
●
builderYou can use this high-bandwidth data to train high-fidelity neural codecs and speech recognition systems.
●
researcherThis dataset offers a massive scale for studying cross-lingual speech representations.