Skip to content

Author

Ibrahim Daoud

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Detecting Voice Conversion Attacks Using Self-Supervised, Handcrafted, and Hybrid Audio Features

Voice conversion attacks are especially challenging to detect because they preserve much of the natural structure of real speech while changing the speaker's identity. In this study, we examine how well handcrafted and self-supervised audio features can detect these attacks, focusing on LFCC and wav2vec 2.0 embeddings, and their score-level fusion. Beyond reporting aggregate performance, our primary contribution is a diagnostic analysis of why performance diverges across different attack types. Experiments were carried out on the ASVspoof 2019 Logical Access dataset and the WaveFake dataset using the same XGBoost-based classification framework for all feature types. Performance was evaluated using Equal Error Rate (EER) and attack-specific miss rates. EER is the primary evaluation metric; perattack miss rate is reported as a secondary diagnostic to highlight variations that may be hidden by the aggregated EER. The results show that LFCC performs better for in-domain vocoder discrimination, while wav2vec 2.0 is more robust on unseen vocoders and performs notably better on voice conversion attacks. The results suggest that self-supervised speech representations capture more robust cues for detecting voice conversion. A back-end sensitivity analysis further shows that the magnitude of this advantage varies substantially with classifier choice, and that some per-attack findings attributed to feature differences are partially classifier-dependent.

Najib Abou Nasr, A. Almutairi, Mohannad Attia et al. · 0 citations