SyncAI.news, a Varaisys broadcasting
SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement
GT

Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen

· 1 min read

ResearcharXiv cs.CL

SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement

arXiv:2609.18009v2 Announce Type: replace-cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.

Original source

This story was published by arXiv cs.CL and written by Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News