
XS
Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang
· 1 min read
ResearcharXiv cs.CV
OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning
arXiv:2609.39150v1 Announce Type: new
Abstract: Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.
Original source
This story was published by arXiv cs.CV and written by Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


