Zero-shot emotional speech synthesis based on feature decoupling and adaptive loss-threshold reweighting
Deep learning-based zero-shot speech synthesis has achieved substantial progress in speaker generalization, but stable modeling remains challenging in fine-grained emotional scenarios. Existing systems often process textual and emotional conditions through shared or closely coupled pathways, which may introduce interfe...