Skip to content
Open access

Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods

Sep 2026 · Journal on Audio, Speech, and Music Processing · Vol 2026 · 0 citations · 19 references

Abstract

Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retraining, making them resource-intensive. Recent target-speaker adaptation techniques fall into two broad categories: (i) fine-tuning-based few-shot methods, which adapt models using small amounts of target-speaker data, and (ii) zero-shot methods, which generalize to unseen speakers without additional training, using a single reference sample during inference. This work compares these two adaptation approaches, along with a training-from-scratch baseline, across different metrics. We show that fine-tuning a non-autoregressive architecture, such as ForwardTacotron, achieves speaker-adaptation quality comparable to zero-shot methods, with lower computational complexity but a lower mean opinion score. As a second contribution, we show that fine-tuning can be made more efficient in terms of data and while largely retaining quality.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.