Target-speaker adaptation in text-to-speech synthesis: a comparison of efficient fine-tuning and zero-shot methods
Neural text-to-speech (TTS) systems can synthesize highly natural speech. A key capability is speaker adaptation, which enables speech synthesis that matches a target speaker’s voice characteristics, such as timbre, pitch, and prosody. Traditional neural approaches require extensive speaker-specific data and full retra...