Aug 2026· Natural Resources for Human Health· 0 citations· 39 references
TL;DR
The results indicate that the combination of facial-region segmentation, latent feature learning and hybrid Transformer architectures enhances the robustness and generalisation capability of deepfake detection significantly.
Abstract
The fast development of deepfake technologies for image generation produces more and more realistic manipulated facial images that are harder to distinguish from the real content. In this paper, we propose a novel Hybrid CNN–Vision Transformer (HybCNNViT) framework for robust deepfake image detection by combining discriminative and generative learning mechanisms. The proposed method presents a new dual-stream architecture that separately encodes inner facial components (eyes, nose and mouth) and outer facial regions with Variational Autoencoder (VAE) and Autoencoder (AE) modules respectively. Also, a hybrid CNN-ViT and CNN-Swin Transformer architecture is used to learn local visual artefacts and global contextual dependencies. In addition, a dedicated preprocessing scheme such as background removal and inner-outer facial object segmentation is further integrated to improve feature relevance and reduce computational complexity. The experimental results on real and fake face datasets under easy, medium and hard manipulation scenarios show the effectiveness of the proposed framework. The proposed Inner–Outer Segment Face Object + HybCNNViT model attains superior accuracy of 86%–88% and outperforms the existing CNN–ViT, ViViT, and CNN–ViTEfficientNet approaches. The results indicate that the combination of facial-region segmentation, latent feature learning and hybrid Transformer architectures enhances the robustness and generalisation capability of deepfake detection significantly.
Deepfake technology, powered by deep learning models, enables the synthesis of highly realistic facial images and videos. However, in recent years, the misuse of deepfakes has posed severe challenges to both individual privacy and social trust. Consequently, this paper systematically reviews research pertaining to deep...
The rapid rise of deepfake technology has raised serious concerns regarding the authenticity of digital media content. This research introduces a hybrid deepfake detection framework that collaboratively combines deep learning and traditional machine learning techniques to improve detection accuracy and robustness. The...
Batini Dhanwanth, Bhargavi Chadalawada, B. Abirami et al.· International Conference Com...· 0 citations
This study develops a video deepfake detection system that addresses the critical challenge of cross-dataset generalization in real-world scenarios. It adopts CLIP (Contrastive Language-Image Pre-training) as the foundation, leveraging its strong vision-language representations for detecting subtle facial manipulations...
Mao-Mao Ling, Włodzimierz Kasprzak· Conference on Computer Scien...· 0 citations
Deep learning has made significant strides recently and has produced amazing lifelike-looking synthetic images called deepfakes, making it very difficult to differentiate between real images and fake images. The abuse of deepfake technology can lead to issues like misinformation, identity fraud, and a loss of confidenc...
Jothi Lakshmi U, B. Sravya, B. M. Reddy et al.· International Conference Com...· 0 citations
A comparative analysis of existing studies is presented to highlight the evolution of deep learning techniques and their effectiveness in improving recognition accuracy and computational efficiency and emerging research directions are outlined to provide insights for future research.
Patel Bhautika Ronak· International journal of res...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.