Skip to content

Band-Attention Modulation Network for Robust Face Forgery Detection

Apr 2024 · 1 citation · 39 references
Computer Science

TL;DR

The Band-Attention Modulation Network (BAM-Net) is proposed, a novel framework that pioneers learnable, fine-grained modulation of frequency components for forgery detection and exhibits exceptional generalization in cross-dataset, cross-compression, and cross-manipulation scenarios.

Abstract

Face forgery detection faces critical challenges in generalizing to unseen manipulation techniques and remaining robust under image compression, which often obscures subtle artifacts. Existing methods typically rely on fixed filters or coarse band separation, lacking the adaptability to learn task-specific spectral cues. To address this, we propose the Band-Attention Modulation Network (BAM-Net), a novel framework that pioneers learnable, fine-grained modulation of frequency components for forgery detection. At its core is the Band-Attention Modulation (BAM) mechanism, which transforms an image into its Discrete Cosine Transform (DCT) spectrogram and learns to dynamically reweight frequency bands along anti-diagonals. This process effectively enhances forgery-related spectral signatures while suppressing less informative ones, simulating an adaptive"inverse compression"that counters information loss. The modulated frequency information is then fused with the spatial domain to guide a lightweight yet effective spatial backbone equipped with distance-decayed attention for comprehensive feature extraction. Extensive experiments on FaceForensics++, Celeb-DF, and DFDC datasets demonstrate that BAM-Net achieves state-of-the-art performance. More importantly, it exhibits exceptional generalization in cross-dataset, cross-compression, and cross-manipulation scenarios, underscoring the vital role of adaptive frequency band modulation in building robust forgery detectors.

View source

Similar papers

Open access Sep 2026

LGF-Net: a spatial-frequency deepfake detection network with learnable gabor filters and gated cross-modal Fusion

LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.

Mukesh Pandey, Deepika Koundal, Sanjeev Kumar · 0 citations
Preprint Sep 2026

Benchmarking Spatial, Spectral, and Self-Supervised Cues for Face Forgery Detection under Realistic Degradation

Face forgery detectors often achieve strong results on controlled benchmarks, but their reliability under realistic image degradations remains limited. This paper presents a standardized benchmark for face forgery detection using the Multi-Dimensional Face Forgery Image (MFFI) dataset and evaluates performance on both...

L. Cunha, Lucas Sotomaior, Lucas Gasperin et al. · 0 citations
Conference Open access 2026

A Comprehensive Investigation on Image-Level Face Forgery Detection in the Spatial Domain

Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deep...

Yu-Jing Zhou · 0 citations
2026

Laplacian Pyramid Reweighting With Progressive Residual Learning for Image Forgery Localization

The increasing realism of image manipulations poses significant challenges for forgery localization. However, existing methods are hindered by the limited adaptability of constrained frequency filters and the dilution of subtle forensic cues in deep networks. To address these challenges, we propose the Laplacian pyrami...

Zhuo-Fei Liu, Wen-Jie Li, Yang Yu · 0 citations
Open access Sep 2026

Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement

A generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement to encourage effective feature disentanglement and improve the discriminability of the learned forgery features is proposed.

Qian Wang, Jia-Qi Feng, Yu Zhou et al. · 0 citations
Open access Oct 2026

Deepfake Image Detection Using Squeeze-and-Excitation Attention and Adaptive Threshold Optimization

Artificial intelligence-generated content (AIGC) has significantly improved the realism of manipulated facial images and videos, creating serious risks for social-network misinformation, content security, and digital trust. This study focuses on visual deepfake image detection rather than multimodal misinformation dete...

Chao-Ran Li · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.