Robust Audio Deepfake Detection Across Controlled and Real-World Datasets Using CNN–Transformers
Abstract
The difference between real and fake speech has been proving difficult due to the development of voice generating technologies. In this research, a hybrid model of convolutional neural networks and the use of Transformer-based sequence modeling to identify audio deepfakes is proposed. The convolutional part obtains spectral representations, and the Transformer obtains time-related dependencies. The model is tested on benchmark and real-life data of a variety of speech samples. The results of the experiment are good; the model achieved 97.33% accuracy, 98.07% precision, 96.57% recall, 97.31% F1-score, 99.65% ROC-AUC, and 2.58% EER. The results indicate that local and global feature modeling are more effeactive to use together and provide more reliability in detection as well as practical application in real-world situations.