Skip to content
Open access

ViT-ConvGDNet: a vision transformer–MobileNet guided decoder network for robust copy-move forgery detection and localization

Sep 2026 · Scientific Reports · 0 citations

Abstract

Copy-move forgery is a common form of image manipulation where a portion of an image is copied and pasted back into the image. This is especially difficult to detect when the forgery has been done on copied areas that have undergone post-processing operations, e.g. rotation, scaling, blurring etc. We propose a new encoder-decoder framework named ViT-ConvGDNet, which integrates the global contextual strengths of Vision Transformers with feature extraction strengths of convolutional operations in MobileNet. Sobel edge detection is added to the encoder to improve the level of awareness of the boundary and sharpness of features. Also, there is Atrous Spatial Pyramid Pooling (ASPP) to obtain multi-scale contextual data that are necessary to accurately perform localization. A layer-wise weighted loss mechanism controls the decoding process, which uses a custom mixture of loss functions at every decoder layer to improve the prediction accuracy. ViT-ConvGDNet makes use of patch-based self-attention mechanisms and is effective at learning long-range dependencies and adapts effectively to images of differing scales and complexity. The performance of the model is better as shown by extensive evaluations on several benchmark datasets such as MICC-F600, MICC-F2000, IMD, Coverage, CoMoFoD, Ardizzone, GRIP, CASIA, and USC-ISI. It has been experimentally demonstrated that ViT-ConvGDNet is more effective than various current deep learning methods and provides a rigorous and scalable solution to problematic copy-move forgery detection and localization.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.