Skip to content
Open access

Vision Transformer Based Digital Image Forgery Detection and Localization Using Global Contextual Feature Learning

Aug 2026 · International Journal of Latest Technology in Engineering Management & Applied Science · 0 citations · 9 references

Abstract

Artificial intelligence has significantly improved digital image editing capabilities, making it increasingly difficult to distinguish authentic images from manipulated ones [5, 7]. This paper proposes a Vision Transformer (ViT)-based framework for digital image forgery detection and localization by leveraging global contextual feature learning [4]. Unlike conventional Convolu-tional Neural Networks (CNNs), Vision Transformers capture long-range dependencies through self-attention mechanisms, enabling more effective identification of manipulated regions [4, 9]. The proposed framework performs image preprocessing, patch extraction, positional encod-ing, transformer-based feature learning, binary classification, and forgery localization. The model is evaluated using publicly available benchmark datasets, including CASIA V2, Co-MoFoD, and FaceForensics++ [20, 48], and its performance is assessed using Accuracy, Pre-cision, Recall, F1-score, Area Under Curve (AUC), Intersection over Union (IoU), and Pixel Accuracy [17, 49]. Experimental results demonstrate that the proposed Vision Transformer framework outperforms conventional CNN-based methods in terms of detection accuracy and localization precision [16, 19]. The proposed approach provides a robust and scalable solution for modern digital image forensics [15] and can be extended to hybrid transformer architectures and video forgery detection in future work.

Read PDF