Skip to content
Conference

A frequency-spatial-temporal deepfake detection framework with parameter-free spatial attention

Sep 2026 · International Conference on Photonic Computing, Algorithms, and Machine Vision · Vol 14320, pp. 143200P - 143200P-12 · 0 citations · 22 references
Engineering

Abstract

Deepfake detection in heavily compressed videos is still challenging because compression often suppresses subtle forgery cues relied upon by existing methods. Existing detectors are observed to rely on frame-level spatial artifacts or computationally expensive spatiotemporal backbones, limiting robustness or efficiency in practical scenarios. FST is proposed as a lightweight video-level framework that jointly exploits frequency-aware cues, native spatial artifacts, and short-term temporal dependency. For each frame, a frequency-aware decomposition-and-reconstruction module performs band-wise DCT filtering and inverse reconstruction to expose manipulation-related spectral irregularities in an imagecompatible representation, while a lightweight spatial branch preserves local RGB artifacts that may be weakened during frequency-aware processing. A Simple, Parameter-Free Attention Module (SimAM) is inserted before global pooling to enhance localized forgery-sensitive responses, and an LSTM-based temporal head aggregates frame-level embeddings for clip-level prediction. Experiments on FaceForensics++ under the heavily compressed C40 setting show that FST achieves 93.10% AUC and 89.50% accuracy with only 19.56M parameters, while also delivering particularly competitive performance on the Deepfakes subset with 96.79% accuracy. Compared with representative image- and video-based detectors, these results indicate that lightweight frequency-spatial-temporal modeling provides a favorable trade-off between detection performance and model complexity for compressed video deepfake detection.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.