DiscoPhon Benchmark for Unsupervised Phoneme Discovery in Speech Models
September 25, 2026
DiscoPhon evaluates unsupervised phoneme discovery using 10 hours of speech across 12 diverse languages. The benchmark tests how well discrete units from models like HuBERT and SpidR map to predefined phoneme inventories via recognition and segmentation metrics.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to evaluate how effectively your unsupervised speech models extract phonemic information.