SyncAI.news, a Varaisys broadcasting
Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition
MK

Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh

· 1 min read

ResearcharXiv cs.CV

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

arXiv:2609.29443v1 Announce Type: new Abstract: Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.

Original source

This story was published by arXiv cs.CV and written by Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News