Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning...
Tianchi Liu, Ze-Yang Song, Tian-Rui Wang et al.· 3 citations
Human evaluation on diagnosed low-quality TTS outputs diagnosed by an AudioLLM shows that LoopTTS can detect perceptually salient errors and correct them with the Refiner, outperforming raw generated audio and practical open-loop re-generation baselines in recovery quality.
Ze-Yang Song, Tianyu Liu, Tian-Rui Wang et al.· 1 citation
EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...
Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.