Skip to content
Conference

Exploiting Hierarchical Representations of Vision Foundation Models for Face Forgery Detection

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 557-564 · 0 citations · 11 references

Abstract

Generalizable face forgery detection has become a critical problem in multimedia forensics as modern face manipulation techniques can generate increasingly realistic facial content. Vision foundation models provide a promising basis for this problem, but existing detectors usually rely on the last-layer visual feature, implicitly assuming that the most transferable forgery evidence is concentrated in the final representation. In a vision transformer, different layers preserve different visual properties, and face manipulation traces may appear as local texture defects, region-level structural conflicts, or high-level semantic inconsistencies. We revisit CLIP from a hierarchical perspective and propose HIAF, a Hierarchical Interaction and Adaptive Fusion framework for face forgery detection. HIAF extracts representations from multiple transformer depths and learns to use them through two dedicated modules. Cross-Layer Feature Interaction (CFI) performs a bottleneck self-attention along the layer dimension, enabling each layer to absorb complementary evidence from other depths and exposing cross-level inconsistencies that are informative for forgery detection. Layer-aware Adaptive Fusion (LAF) then assigns dynamic weights to different layers, adaptively promoting the layers whose forensic cues are most reliable while suppressing less discriminative ones. To further stabilize the layer-wise feature spaces, we introduce a multi-view contrastive objective that treats each layer as an independent view and regularizes real samples into compact manifolds. Extensive experiments show that HIAF achieves an average AUC of 88.93% in cross-dataset evaluation and 96.54% in cross-manipulation evaluation, outperforming recent state-of-the-art methods.

View source

Similar papers

Preprint Aug 2026

Learning Unified Video and Image Representation for Video Face Forgery Detection

A novel framework, UVIF, that utilizes additional annotated images to provide fine-grained supervision for detecting partial forgeries in videos, which outperforms state-of-theart methods in detecting partially forged videos while introducing no additional computational overhead is proposed.

Haotian Liu, Y. Liu, Guoying Zhao et al. · 0 citations
Conference Open access 2026

A Comprehensive Investigation on Image-Level Face Forgery Detection in the Spatial Domain

Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deep...

Yu-Jing Zhou · 0 citations
Preprint Sep 2026

DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection

As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulation...

Feng-Ming Gu, Ming-Jie He, Zong-Hui Guo et al. · 0 citations
Preprint Sep 2026

ForensicZoom: Adaptive Visual Inspection with Multimodal LLMs for Industrial-Grade Face Forgery Detection

Reliable face forgery detection is critical to the security of online identity verification systems, where missed attacks compromise security and excessive false positives disrupt legitimate users. Specialized forensic detectors achieve strong detection performance but provide limited interpretability, while multimodal...

Hang Zhou, Yi-Ming Tang, Kun Yu et al. · 0 citations
Conference Aug 2026

Explainable Deep Learning Framework for Accurate Detection and Interpretation of Copy-Move Forgery in Digital Images

In the era of advanced digital image generation techniques such as copy-move forgery, the NDV is having increasing concerns regarding image authenticity particularly with the spread of digital images through social media, media and courts. This type of forgery, where a part of an image is replicated and pasted on the s...

Shaheena K. V., D. S · 0 citations
Aug 2026

End-to-end Multi-face Forgery Detection via Progressively-explicit Queries

This paper proposes an end-to-end Transformer-based framework, termed Progressively Explicit Query Network (PEQNet), for multi-face forgery detection and localization, and introduces triple contrastive learning to model the mutual exclusivity among real, fake, and background regions.

Peng-Wen Dai, Xiaomeng Wen, Feiyang He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.