SyncAI.news, a Varaisys broadcasting
Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection
XL

Xinzhe Li, Youzhi Tu, Kong Aik Lee

· 1 min read

ResearcharXiv cs.AI

Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

arXiv:2609.33394v1 Announce Type: new Abstract: The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.

Original source

This story was published by arXiv cs.AI and written by Xinzhe Li, Youzhi Tu, Kong Aik Lee. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News