Skip to content
Open access

Cross-Domain Generalization From Synthesized to Human-Imitated Speech Detection

2026 · IEEE Access · Vol 14, pp. 147755-147768 · 0 citations · 35 references

Abstract

The increasing prevalence of spoofing attacks in automatic speaker verification systems underscores the critical need for robust detection methods to ensure security and reliability. Although current spoof speech datasets primarily focus on synthesized speech and voice conversion attacks, human-imitated speech remains relatively underexplored. To address this gap, this paper proposes a human-imitated speech dataset designed for machine learning models and acoustic feature representations and evaluate various deep-learning approaches to improve existing spoofing-countermeasure systems. These systems are then evaluated using both synthesized and the proposed human-imitated speech datasets, focusing on models trained on the ASVspoof 2019 LA dataset. The model performs well when trained and tested on synthesized speech, but it performs significantly worse when tested on the proposed human-imitated speech dataset. When trained on the proposed imitated speech dataset, the model detects imitated speech notably better. The results show the importance of the proposed approaches in developing more robust spoofing detection methods.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.