Skip to content
Open access

Speech-GAN: A Black-Box Generative Adversarial Network Attack against Automatic Speech Recognition systems

Jul 2026 · Vietnam Journal of Computer Science · 0 citations

Abstract

Automatic speech recognition (ASR) systems rely on deep learning models to transform human speech into text and actionable commands. Despite their effectiveness, such systems are vulnerable to adversarial audio, which can cause incorrect transcriptions and lead to unintended system behavior. Understanding these vulnerabilities is therefore essential for the safe deployment of ASR in security-sensitive and safetycritical contexts. In this paper, we present Speech-GAN, a generative adversarial network designed to perform black-box untargeted attacks against deep learning-based ASR systems by generating adversarial audio samples.We validate the effectiveness of Speech-GAN through 10,000 attack runs conducted on 1,000 audio samples spanning 10 command words, targeting theWav2Vec 2.0 ASR model. Speech-GAN achieves success rates exceeding 99%, with average signal-to-noise ratios of -9.53 dB for basic attacks and -11.21 dB for semantically constrained attacks. A set of 10 humans confirmed the effectiveness of Speech-GAN: they understand 91.4% (resp. 97.6%) of the semantically constrained (resp. basic) adversarial audios as the intended command.

Read PDF