LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
Abstract
Deepfakes have become increasingly realistic due to recent advances in face manipulation techniques, making reliable detection in unconstrained environments more challenging. Existing spatial-frequency deepfake detection methods often rely on fixed hand-crafted frequency transforms and simple fusion strategies, which may limit their adaptability and cross-dataset generalization. To address these limitations, we propose LGF-Net, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection. The Frequency Representation Module employs learnable Gabor filters and a frequency-aware attention mechanism to capture manipulation-specific spectral patterns. Moreover, the Spatial Representation Module uses multi-rate dilated convolutions to model both subtle local artifacts and long-range structural inconsistencies, while a gated cross-modal fusion module integrates the two representations into a compact forensic descriptor. Experimental results on FF++ (HQ), Celeb-DF (V2), DPDC, and DFD show that LGF-Net achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
Deepfake detection in heavily compressed videos is still challenging because compression often suppresses subtle forgery cues relied upon by existing methods. Existing detectors are observed to rely on frame-level spatial artifacts or computationally expensive spatiotemporal backbones, limiting robustness or efficiency...
Yi-Zhi Wang· International Conference on...· 0 citations
Real-time and accurate dense small-scale face detection constitutes a critical technology in computer vision. However, in complex environments, this task presents significant challenges, including high target density, minute scales, frequent mutual occlusions, and background noise interference. These factors often caus...
Haomin Li· Poster Volume 0008 The 2026...· 0 citations
Face Super-Resolution (FSR) faces a critical challenge in balancing reconstruction quality with computational efficiency when deployed on resource-constrained devices. While existing methods leverage CNNs or transformers, they are constrained by limited receptive fields or high computational complexity. To achieve both...
Yun-Zhe Xu, Chen Zhao, Zhi-Zhou Chen et al.· IEEE Transactions on Image P...· 0 citations
FAViT (Frequency-Aware Vision Transformer), a hybrid architecture capable of jointly utilizing spatial- and frequency-domain forensic information by the means of a bidirectional cross-attention fusion scheme, is presented.
Wasin Alkishri, Shahid Kamal, Jabar H. Yousif· Information· 0 citations
Infrared and visible image fusion aims to integrate complementary information from different modalities to generate images with both salient targets and rich textures. However, existing methods mainly rely on spatial feature modeling and lack an explicit mechanism to exploit frequency-aware representations, limiting th...
Yufeng Li, Lei Yu, Chuan-Long Xie et al.· Engineering Research Express· 0 citations
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible s...
Nasir Saleem, Ahmad Ali, Zhuo-Qi Zeng et al.· International Journal of Int...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.