SSE: Multimodal Generative Model for User-Guided Audio Remixing
September 25, 2026
Spot, Separate, and Enhance (SSE) is a multimodal generative model that rebalances audio, removes sources, and reduces reverberation using video and text prompts. The approach outperforms reconstruction-based baselines by focusing on controllability and creative remixing quality.
HOW THIS AFFECTS YOU
●
builderThe DegradedMix dataset and SSE architecture provide a new foundation for generative audio enhancement tools.
●
designerYou can use text and video as intuitive controls for complex audio post-production tasks.