Aug 2026· Journal of Forensic Sciences· 0 citations· 9 references
Medicine
TL;DR
EMBNet is proposed, a task-oriented deepfake speech detection framework that integrates efficient channel attention (ECA) with a multi-scale bottleneck to enhance the representation of subtle acoustic anomalies and may support forensic audio authenticity assessment by improving the discrimination between genuine and manipulated speech.
Abstract
With the rapid advancement of generative artificial intelligence, deepfake speech has emerged as a significant threat to digital audio authenticity, posing challenges for forensic analysis and legal applications. In this study, we propose EMBNet, a task-oriented deepfake speech detection framework that integrates efficient channel attention (ECA) with a multi-scale bottleneck to enhance the representation of subtle acoustic anomalies. The ECA module adaptively emphasizes critical feature channels, while the multi-scale bottleneck captures local and hierarchical spoofing traces across multiple time-frequency resolutions. The proposed framework effectively balances fine-grained local detail modeling with global hierarchical representation, improving the detection of weakly manifested spoofing patterns. Extensive experiments on the ASVspoof 2019 Logical Access dataset demonstrate that EMBNet significantly outperforms existing baseline and state-of-the-art methods, achieving an EER of 2.67%, an AUC of 97.32%, and an F1-score of 97.28%. Ablation studies further confirm the complementary contributions of the ECA and multi-scale modules to overall performance. The proposed approach demonstrates promising performance for forensic audio analysis and may support forensic audio authenticity assessment by improving the discrimination between genuine and manipulated speech.
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
The results indicate that multi-domain feature fusion offers a practical and generalizable solution for real-world deepfake audio detection, particularly in environments involving diverse codecs and synthesis techniques.
Akshat Chhatriwala, Ishita Akolkar, Namrata Shroff et al.· ITEGAM- Journal of Engineeri...· 0 citations
Recent advances in generative artificial intelligence have enabled highly realistic speech synthesis using text-to-speech (TTS), voice conversion (VC), and neural voice cloning techniques, posing significant security threats to Automatic Speaker Verification (ASV) systems. Conventional handcrafted features such as Mel-...
K. S. Kumar, Madduluri Suneetha, K. R. Anudeep Laxmi Kanth et al.· International journal of com...· 0 citations
Recent advances in text-to-audio (TTA) and audio-to-audio (ATA) generation models have enabled the creation of highly realistic environmental sounds, raising growing concerns about malicious audio manipulation in real-world scenarios. To address this emerging threat, the ESDD 2026 Challenge was introduced as the first...
Sanghyeok Chung, Seungsang Oh, Donggun Kim et al.· Proceedings of the Thirty-Fi...· 0 citations
The proliferation of sophisticated audio deepfake technology poses a significant threat to digital voice authentication and forensic verification systems. This research addresses this challenge by developing and evaluating a lightweight hybrid Convolutional Neural Network and Recurrent Neural Network (CNN-RNN) architec...
Muh. Hajar Akbar, Nurfitria Ningsi, Aldi et al.· Journal of Information Syste...· 0 citations
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD un...
Si-Qing Qin, Zhe Li, K. Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.