A generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement to encourage effective feature disentanglement and improve the discriminability of the learned forgery features is proposed.
Abstract
Existing deepfake detectors often perform well on in-domain data but generalize poorly to unseen datasets or manipulation methods. This limitation is largely attributed to their reliance on dataset-specific semantic cues rather than transferable forgery patterns. To address this limitation, we propose a generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement. A Phase-Amplitude Frequency Enhancement (PAFE) module enhances subtle spectral artifacts introduced during deepfake generation. We then feed the enhanced representations into an asymmetric dual-branch architecture that separates content-related information from forgery-related features. The content branch models facial semantics, while the forgery branch extracts discriminative forgery features with reduced content interference. A spatial self-attention module further refines the forgery features. We optimize the framework using image-level reconstruction loss, feature-level contrastive loss, and classification loss. Together, these objectives encourage effective feature disentanglement and improve the discriminability of the learned forgery features. Extensive experiments on several widely used deepfake benchmarks show that the proposed framework achieves competitive detection performance and improved cross-domain generalization compared with existing methods.
This work explores an approach that integrates wavelet-based frequency analysis with deep learning to enhance deepfake detection, and suggests that wavelet sub-bands expose manipulation cues that are useful for detecting unseen fake classes, but they should not be interpreted as a uniform robustness improvement.
LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
Deepfake detection in heavily compressed videos is still challenging because compression often suppresses subtle forgery cues relied upon by existing methods. Existing detectors are observed to rely on frame-level spatial artifacts or computationally expensive spatiotemporal backbones, limiting robustness or efficiency...
Yi-Zhi Wang· International Conference on...· 0 citations
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
This work proposes an innovative Environment-Invariant Subspace Learning (EISL) framework, which aims to disentangle features into orthogonal forgery-relevant invariant factors and environment-related residual factors via a learnable low-rank projection and designs an Environmental Intervention module that generates di...
Sheng-Hao Chen, Hao Jia, Chen Li et al.· 0 citations
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulation...
Feng-Ming Gu, Ming-Jie He, Zong-Hui Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.