Skip to content
Preprint

Your Voice Cloning System is Secretly a Voice Anonymizer

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work proposes repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining, and introduces an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility.

Abstract

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER $\approx$ 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.

View source

Similar papers

Preprint Sep 2026

VoxTubeS: Distributable Speaker-Anonymized Synthetic Speech Corpora and Their Analysis

VoxTubeS is presented, a family of speaker-anonymized synthetic speech corpora designed for redistribution, comprising three method families and seven variants derived from the VoxTube corpus, which is distributed under CC BY-NC-SA 4.0.

Zhe Zhang, Ye-Xin Lu, Junichi Yamagishi · 0 citations
Preprint Sep 2026

Decaf: A privacy preserving speech codec using speaker disentanglement and canonical voice conversion

We present DECAF, a privacy preserving neural speech codec that obfuscates a speaker's voice while preserving linguistic content while maintaining automatic speech recognition (ASR) performance at very low bitrates, inspired by decaffeination. At the transmitter end, speech is encoded into speaker independent content e...

M. Siam, D. Sharma, S. Kruchinin et al. · 0 citations
Preprint Sep 2026

Reducing Speaker Residual by Considering Pinhole Effect in Voice Anonymization

Voice anonymization aims to protect privacy by suppressing speaker identity while preserving linguistic content and prosody. However, residual speaker attributes in non-identity representations may still increase linkability and weaken privacy protection. To this end, this paper proposes a fine-tuning strategy with a p...

Ze-Yan Liu, Wei Jiang, Li-Ping Chen et al. · 0 citations
Preprint Aug 2026

DP-VOXLET: Provable Speaker Anonymization for Disentangled Speech Representations

Systems for speaker anonymization obfuscate the speaker of an utterance, while maintaining its original semantic contents and prosody. Recent solutions for speaker anonymization rely on learned representations that disentangle an utterance into semantic contents and speaker properties. To anonymize an utterance, these...

Ivoline C. Ngong, Jack D'Iorio, Hailey Schoppe et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech

Recent zero-shot text-to-speech (ZS-TTS) systems can reproduce a speaker's voice with high fidelity from only a few seconds of reference speech, raising concerns over unauthorized voice cloning and impersonation. Speaker identity unlearning has recently emerged as an approach to selectively suppress this capability for...

Hyoeun Kim, Y. Lee, Kyuhong Shim · 0 citations
#artificial intelligence Preprint Sep 2026

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (...

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.