
CB
Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene
· 1 min read
ResearcharXiv cs.CV
Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling
arXiv:2610.03380v1 Announce Type: new
Abstract: Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.
Original source
This story was published by arXiv cs.CV and written by Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


