Aug 2026· International Conference Computational Vision and Bio Inspired Computing· pp. 1-9· 0 citations· 18 references
Abstract
The recent upsurge in the development of sophisticated generative models has significantly improved the visual realism and semantic coherence of synthetic images, thus presenting a major challenge to the field of multimedia forensics. The conventional approaches often rely on either artifacts or high level semantic cues, limiting their robustness when handling images/videos generated by models that were previously unseen. The proposed work addresses this problem by developing a novel forensic consistency learning framework that is conditioned on semantic content. Specifically, the model leverages pretrained DINO (self-DIstillation with NO labels) encoder for visual content, forensic features capturing latent acquisition characteristics and image-level CLIP (Contrastive Language-Image Pretraining) features for global semantic content. A forensic predictor module estimates the expected forensic features conditioned on semantic information, enabling capture of inconsistencies between visual content and underlying artifacts. Additionally, a patch-level anomaly score module enables robust image-level prediction. The method was evaluated under cross-generator setting, by training on GenImage dataset augmented with ProGAN, and evaluated on UniversalFakeDetect (UFD) benchmark. The model achieves 91.5% Area Under ROC Curve, 92.2% Average Precision and 84.1% classification accuracy, significantly outperforming existing baselines like UFD and ResNet (upto 6-10% improvement). Extensive ablations further validate the effectiveness of proposed semantic-conditioned forensic modeling for open-world image authentication.
A generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries is proposed, and Sparse-Constraint Rectified Flow is introduced, a detector-oriented adaptation of F...
Jiangling Zhang, Shuxuan Gao, Zeyu Chen et al.· 0 citations
This work proposes a compositional forensic visual prompt learning framework that operates entirely in the visual feature space and employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches.
Fangling Jiang, Qi Li, Bing Liu et al.· 0 citations
Steganography using deep learning can preserve pixel-level image quality while still changing object, attribute, or relational information in captions generated by vision-language models (VLMs). This caption drift creates a detection channel that is not measured by global image-embedding similarity alone. This paper pr...
This paper hypothesize, and empirically demonstrate, that the intrinsic world understanding of JEPA models can be used as a strong prior for a deepfake detector, and proposes MoE-JEPA, a dual-stream architecture for deepfake detection.
FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token.
Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska· 0 citations
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task,...
Javier Muñoz-Haro, Ruben Tolosana, Rubén Vera-Rodríguez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.