Robust Audio Deepfake Detection Across Controlled and Real-World Datasets Using CNN–Transformers
The difference between real and fake speech has been proving difficult due to the development of voice generating technologies. In this research, a hybrid model of convolutional neural networks and the use of Transformer-based sequence modeling to identify audio deepfakes is proposed. The convolutional part obtains spe...