SyncAI.news, a Varaisys broadcasting
ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
JC

Jiaran Cai, Xingpei Ma, Shenneng Huang

· 1 min read

ResearcharXiv cs.CV

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

arXiv:2609.29225v1 Announce Type: new Abstract: Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.

Original source

This story was published by arXiv cs.CV and written by Jiaran Cai, Xingpei Ma, Shenneng Huang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News