
IV
Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu
· 1 min read
ResearcharXiv cs.AI
Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
arXiv:2609.29169v1 Announce Type: cross
Abstract: We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site
Original source
This story was published by arXiv cs.AI and written by Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


