Xiaomi-CocktailASR-1: LLM-based Target Speaker ASR
September 11, 2026
Xiaomi-CocktailASR-1 is an end-to-end architecture that uses reference speech as voiceprint prompts to transcribe target speakers in multi-speaker environments. It achieves single-speaker performance comparable to mainstream models and includes a negative sample rejection capability.
HOW THIS AFFECTS YOU
●
builderThis enables more robust voice interface performance in noisy, multi-person settings.
●
researcherThe use of voiceprint prompts in LLM-based ASR avoids the need for traditional explicit speech separation.