Skip to content
Preprint

Unifying Semantic Priors and High-Frequency Traces: Enhancing V-JEPA with Mixture-of-Experts for Robust Synthetic Image Forensics

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

This paper hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector, and proposes MoE-JEPA, a dual-stream architecture for deepfake detection.

Abstract

The unchecked proliferation of manipulated images on social media platforms has increased the spread of misinformation, posing a severe threat to public trust and information integrity. Modern deepfake detectors typically rely on Vision Transformers (ViTs) to capture the low-level inconsistencies that characterize fully synthetic or locally tampered images. However, the global understanding of such foundation models is not enough to discriminate alone between real and fake multimedia content, especially in challenging scenarios where images are compressed or transmitted through social media. In this paper we pioneer the application of Joint-Embedding Predictive Architecture (JEPA) models to deepfake detection, taking advantage of the generalized representation of visual reality that such World Models have exhibited. We hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector. To fully exploit JEPA capabilities, we propose MoE-JEPA, a dual-stream architecture for deepfake detection. By enhancing a V-JEPA 2 backbone with a Residual Mixture-of-Experts (MoE) mechanism, along with a noise stream branch, our model dynamically internalizes forensic knowledge. Furthermore, a Gated Attention Multiple Instance Learning (MIL) module is employed to ensure precise spatial semantic understanding. Evaluated on the SID-Set benchmark, comprising 300K AI-generated, tampered and authentic images, MoE-JEPA establishes a new state-of-the-art with an accuracy of 95.54%, successfully outperforming vastly larger models.

View source

Similar papers

Conference Aug 2026

Semantic-Conditioned Forensic Consistency Model for Open-World Image Authenticity Verification

The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cue...

Venkata Satya Renuka Devi Bhamidipati, Srinivasa Rao Chanamallu, Sudheer Gopinathan · 0 citations
2026

MF2DA: Multi-Level Feature Fusion for Robust Detection and Attribution of Universal AI-Generated Images

The proliferation of hyper-realistic AI-generated images poses significant threats to digital information integrity and forensic accountability. Existing detection methodologies, however, face three critical bottlenecks: vulnerability to real-world distortions such as social media compression, inadequate specialization...

Wen-Peng Mu, Qiang Xu, Yi-Ning Zhang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitt...

Xue-Chao Zou, Yi Zhou, Kai Li et al. · 0 citations
Preprint Aug 2026

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

This work proposes a compositional forensic visual prompt learning framework that operates entirely in the visual feature space and employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches.

Fangling Jiang, Qi Li, Bing Liu et al. · 0 citations
Conference Open access 2026

Passive Detection of GAN-Generated Images: A Structured Investigation of Spatial and Frequency-Domain Approaches

The rapid advancements in generative adversarial networks (GANs) have led to the production of highly realistic synthetic images, posing severe threats to the credibility and authenticity of digital media across social platforms, news outlets, and official documents. Passive detection methods tackle this problem by ide...

Lin-Jie Lu · 0 citations
Preprint Sep 2026

A Multi-View and Confusion-Guided Ensemble Framework for Robust Synthetic Image Attribution

This report presents a multi-view and confusion-guided ensemble framework for the Synthetic Image Attribution Challenge of the DLMMDD Workshop at ICANN 2026, and introduces a dedicated binary expert classifier that is selectively activated under low-confidence conditions.

Zuo-Min Qu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.