A frequency-spatial-temporal deepfake detection framework with parameter-free spatial attention
Abstract
Deepfake detection in heavily compressed videos is still challenging because compression often suppresses subtle forgery cues relied upon by existing methods. Existing detectors are observed to rely on frame-level spatial artifacts or computationally expensive spatiotemporal backbones, limiting robustness or efficiency in practical scenarios. FST is proposed as a lightweight video-level framework that jointly exploits frequency-aware cues, native spatial artifacts, and short-term temporal dependency. For each frame, a frequency-aware decomposition-and-reconstruction module performs band-wise DCT filtering and inverse reconstruction to expose manipulation-related spectral irregularities in an imagecompatible representation, while a lightweight spatial branch preserves local RGB artifacts that may be weakened during frequency-aware processing. A Simple, Parameter-Free Attention Module (SimAM) is inserted before global pooling to enhance localized forgery-sensitive responses, and an LSTM-based temporal head aggregates frame-level embeddings for clip-level prediction. Experiments on FaceForensics++ under the heavily compressed C40 setting show that FST achieves 93.10% AUC and 89.50% accuracy with only 19.56M parameters, while also delivering particularly competitive performance on the Deepfakes subset with 96.79% accuracy. Compared with representative image- and video-based detectors, these results indicate that lightweight frequency-spatial-temporal modeling provides a favorable trade-off between detection performance and model complexity for compressed video deepfake detection.