A survey of threats against voice authentication and anti-spoofing systems
Abstract
Voice authentication has evolved from traditional systems based on hand-engineered acoustic features to deep learning models that learn speaker representations directly from audio. This progress has improved accuracy, but it has also opened new attack surfaces. Prior surveys treat these threats in isolation; we instead provide a unified analysis of the four primary attack vectors against voice authentication: data poisoning, adversarial perturbations, audio deepfakes, and adversarial spoofing. We organize them along three axes: what the attacker knows, how the attack is delivered, and how broadly it generalizes across speakers and inputs. Comparing these vectors reveals a clear gap in threat maturity. Adversarial perturbations are effective when the attacker has full model access, but often fail once the audio is played over the air. Deepfakes, by contrast, scale easily, succeed without model access, and routinely fool both machines and human listeners. Existing Anti-Spoofing Countermeasures (CMs) remain largely reactive and can themselves be bypassed by adaptive attacks. We close by mapping open challenges and arguing for standardized threat models and defenses that act proactively, in real time.