SyncAI.news, a Varaisys broadcasting
When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation
ZX

Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng

· 1 min read

ResearcharXiv cs.CV

When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

arXiv:2609.36685v1 Announce Type: new Abstract: Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.

Original source

This story was published by arXiv cs.CV and written by Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu, Ming Meng. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News