
GT
Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
· 1 min read
ResearcharXiv cs.CL
SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement
arXiv:2609.18009v2 Announce Type: replace-cross
Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.
Original source
This story was published by arXiv cs.CL and written by Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


