Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods
Abstract
Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retraining, making them resource-intensive. Recent target-speaker adaptation techniques fall into two broad categories: (i) fine-tuning-based few-shot methods, which adapt models using small amounts of target-speaker data, and (ii) zero-shot methods, which generalize to unseen speakers without additional training, using a single reference sample during inference. This work compares these two adaptation approaches, along with a training-from-scratch baseline, across different metrics. We show that fine-tuning a non-autoregressive architecture, such as ForwardTacotron, achieves speaker-adaptation quality comparable to zero-shot methods, with lower computational complexity but a lower mean opinion score. As a second contribution, we show that fine-tuning can be made more efficient in terms of data and while largely retaining quality.