This paper hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector, and proposes MoE-JEPA, a dual-stream architecture for deepfake detection.
Abstract
The unchecked proliferation of manipulated images on social media platforms has increased the spread of misinformation, posing a severe threat to public trust and information integrity. Modern deepfake detectors typically rely on Vision Transformers (ViTs) to capture the low-level inconsistencies that characterize fully synthetic or locally tampered images. However, the global understanding of such foundation models is not enough to discriminate alone between real and fake multimedia content, especially in challenging scenarios where images are compressed or transmitted through social media. In this paper we pioneer the application of Joint-Embedding Predictive Architecture (JEPA) models to deepfake detection, taking advantage of the generalized representation of visual reality that such World Models have exhibited. We hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector. To fully exploit JEPA capabilities, we propose MoE-JEPA, a dual-stream architecture for deepfake detection. By enhancing a V-JEPA 2 backbone with a Residual Mixture-of-Experts (MoE) mechanism, along with a noise stream branch, our model dynamically internalizes forensic knowledge. Furthermore, a Gated Attention Multiple Instance Learning (MIL) module is employed to ensure precise spatial semantic understanding. Evaluated on the SID-Set benchmark, comprising 300K AI-generated, tampered and authentic images, MoE-JEPA establishes a new state-of-the-art with an accuracy of 95.54%, successfully outperforming vastly larger models.
The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cue...
The proliferation of hyper-realistic AI-generated images poses significant threats to digital information integrity and forensic accountability. Existing detection methodologies, however, face three critical bottlenecks: vulnerability to real-world distortions such as social media compression, inadequate specialization...
Wen-Peng Mu, Qiang Xu, Yi-Ning Zhang et al.· IEEE Transactions on Informa...· 1 citation
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitt...
This work proposes a compositional forensic visual prompt learning framework that operates entirely in the visual feature space and employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches.
Fangling Jiang, Qi Li, Bing Liu et al.· 0 citations
The rapid advancements in generative adversarial networks (GANs) have led to the production of highly realistic synthetic images, posing severe threats to the credibility and authenticity of digital media across social platforms, news outlets, and official documents. Passive detection methods tackle this problem by ide...
This report presents a multi-view and confusion-guided ensemble framework for the Synthetic Image Attribution Challenge of the DLMMDD Workshop at ICANN 2026, and introduces a dedicated binary expert classifier that is selectively activated under low-confidence conditions.
Zuo-Min Qu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.