
PW
Pu Wang, Hugo Van hamme
· 1 min read
ResearcharXiv cs.CV
ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion
arXiv:2609.26648v2 Announce Type: replace-cross
Abstract: Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams together with modality-agnostic query tokens, enabling direct interaction among available modality inputs. Modality dropout further improves robustness when input streams are unavailable. ROAM-ASD achieves state-of-the-art performance across five ASD benchmarks: 98.8% mAP on WASD, 87.9% on UniTalk, 96.5% on AVA, 99.3% on ASW, and 98.2% on Talkies, improving over previous best systems by 5.1, 4.7, 0.9, 1.0, and 2.1 mAP points, respectively. ROAM-ASD also substantially improves zero-shot cross-dataset generalization and remains robust to missing observations.
Original source
This story was published by arXiv cs.CV and written by Pu Wang, Hugo Van hamme. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


