Skip to content

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Jul 2026 · arXiv.org · Vol abs/2607.16870 · 0 citations · 40 references
Computer Science

TL;DR

This work introduces Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and proposes SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder.

Abstract

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.

View source

Similar papers

Preprint Oct 2026

Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs

Encoder-based speech large language models (Speech-LLMs) commonly employ pretrained speech encoders that prioritize linguistic content but may discard fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding. Encoder-free Speech-LLMs instead map Mel-spectrogram features directly...

Mo-Han Shi, Ru-Chao Fan, S. Sivasankaran et al. · 0 citations
Preprint Aug 2026

SURE-Voice: A Front-End Baseline for Speech-Evidence Filtering in Speech LLMs

Speech language models (speech LLMs) can generate plausible outputs from audio that contains no usable speech evidence. We study this failure as a pre-generation support-estimation problem and present SURE-Voice, a training-free front end that decides whether an audio prompt contains intelligible speech evidence before...

Mengzhe Geng · 0 citations
Preprint Aug 2026

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

A multilingual zero-shot TTS system that supports 14 languages and performs both intra-lingual and cross-lingual reference cloning without requiring transcripts of audio prompts, which achieves the best average overall rank in human evaluation among recent open-source and commercial systems.

Hua-Xuan Wang, Hui-Min Wang, Ruiyu Zhang et al. · 0 citations
Preprint Aug 2026

Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages

Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.

Varun Rai, J. PavanKumar, Sujith Pulikodan et al. · 0 citations
#natural language process... Preprint Sep 2026

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, these models struggle with predicting high-bitrate speech tokens in LLMs, and face the challenge of relying on S2ST training data with ideall...

Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.