Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
This work introduces Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and proposes SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder.