SyncAI.news, a Varaisys broadcasting
Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
JJ

Jihoo Jung, Youngjoon Jang, Joon Son Chung

· 1 min read

ResearcharXiv cs.CV

Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs

arXiv:2609.31193v1 Announce Type: new Abstract: Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.

Original source

This story was published by arXiv cs.CV and written by Jihoo Jung, Youngjoon Jang, Joon Son Chung. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News