SyncAI.news, a Varaisys broadcasting
Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval
XC

Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei

· 1 min read

ResearcharXiv cs.CV

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

arXiv:2604.23692v2 Announce Type: replace-cross Abstract: Audio-driven facial animation is essential for immersive digital interaction, yet existing frameworks struggle to reconcile real-time streaming with high-fidelity personalization. Current methods either rely on latency-inducing audio look-ahead, or ask users to record scripted calibration sequences to pre-encode static identity embeddings that fail to capture dynamic idiosyncrasies. We present an end-to-end framework for personalized audio-driven facial motion generation, supporting causal, zero-lookahead streaming. We introduce two key innovations: (1) a causal multi-resolution motion tokenizer that captures both global temporal context and high-frequency articulatory details, and (2) a multi-modal style retriever that extracts stylistic priors from unstructured reference libraries by jointly querying ongoing audio and motion. Unlike prior retrieval mechanisms restricted to curated, fixed-size, or audio-only style banks, our design accepts arbitrary footage of the target identity, enabling personalization from a handful of casually recorded clips. By integrating these components, our method outperforms state-of-the-art approaches in lip-sync accuracy, identity consistency, and perceived realism, while preserving the streaming constraints of real-time telepresence. Code is available at https://github.com/xg-chu/Fallingwater.

Original source

This story was published by arXiv cs.CV and written by Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News