[arXiv]score: 0.24
Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge
August 17, 2026
VibeVoice-ASR-7B improves speaker diarization performance by reducing tcpMER from 18.30% to 16.73% using random leading-silence cropping and exponential moving average training. For conversational speech understanding, Qwen3-Omni-30B-A3B-Instruct is fine-tuned on synthetic question-answer pairs generated through multimodal candidate generation and silent-audio filtering.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy