Skip to content
Open access

A Multi-Domain Feature Framework for Robust Deepfake Audio Detection

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · 0 citations

TL;DR

The results indicate that multi-domain feature fusion offers a practical and generalizable solution for real-world deepfake audio detection, particularly in environments involving diverse codecs and synthesis techniques.

Abstract

Audio deepfakes generated by modern text-to-speech and voice conversion systems pose serious threats to security, privacy, and trust in digital communication. This study proposes a multi-domain feature fusion framework for robust deepfake audio detection under realistic, in-the-wild conditions. A large-scale dataset comprising 31,780 audio samples, evenly split between genuine and synthetic speech and covering diverse speakers, languages, and recording environments, is utilized. Acoustic, compression-related, emotional, phase-based, prosodic, and statistical–spectral features are extracted and fused, and classification is performed using a lightweight fully connected neural network evaluated via stratified five-fold cross-validation. The proposed system achieves an average validation accuracy of 98.03% and an AUC of 0.998, demonstrating strong and stable discriminative performance. Ablation experiments and SHAP-based analysis highlight the critical role of compression-related features in enhancing robustness when combined with other feature domains. While the study focuses on audio-only detection and does not address adversarial or multimodal scenarios, the results indicate that multi-domain feature fusion offers a practical and generalizable solution for real-world deepfake audio detection, particularly in environments involving diverse codecs and synthesis techniques.

Read PDF

Similar papers

Open access Aug 2026

EMBNet: Multi-scale feature learning with efficient channel attention for deepfake speech detection.

EMBNet is proposed, a task-oriented deepfake speech detection framework that integrates efficient channel attention (ECA) with a multi-scale bottleneck to enhance the representation of subtle acoustic anomalies and may support forensic audio authenticity assessment by improving the discrimination between genuine and ma...

Haitao Yang, Fen Li, Xin Cai et al. · 0 citations
Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations
Open access Aug 2026

A Hybrid MFCC–WavLM Feature Fusion for Audio Deepfake Detection

Recent advances in generative artificial intelligence have enabled highly realistic speech synthesis using text-to-speech (TTS), voice conversion (VC), and neural voice cloning techniques, posing significant security threats to Automatic Speaker Verification (ASV) systems. Conventional handcrafted features such as Mel-...

K. S. Kumar, Madduluri Suneetha, K. R. Anudeep Laxmi Kanth et al. · 0 citations
Open access Sep 2026

Audio Deepfake Detection Using Dual-Branch CNN with Shared Weights

The detection of audio deepfakes has emerged as a significant problem in the field of voice biometrics systems, aiming to distinguish real human voices from those generated by Artificial Intelligence (AI). With synthetic voice becoming increasingly high-quality, it is more likely that such a voice will be abused for il...

Zainab A. Jawad, Ahmed J. Obaid · 0 citations
Preprint Aug 2026

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the...

Cunhang Fan, Jun-Qin Cao, Tian Gao et al. · 0 citations
Open access Sep 2026

Robust Audio Deepfake Detection Across Controlled and Real-World Datasets Using CNN–Transformers

The difference between real and fake speech has been proving difficult due to the development of voice generating technologies. In this research, a hybrid model of convolutional neural networks and the use of Transformer-based sequence modeling to identify audio deepfakes is proposed. The convolutional part obtains spe...

Rafal Akeel, Belal Al-Khateeb · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.