
JR
Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski
· 1 min read
ResearcharXiv cs.CV
DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline
arXiv:2609.23830v1 Announce Type: new
Abstract: Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.
Original source
This story was published by arXiv cs.CV and written by Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


