This work proposes AuEmoChat, a CSS framework for authentic emotion understanding and rendering, and develops AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories.
Abstract
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. Furthermore, we propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech. The code and speech demos will be available at: https://github.com/AI-S2-Lab/AuEmoChat.
ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions, is introduced.
Arthur Peuvot, Romaric Besançon, Gaël de Chalendar et al.· 0 citations
EmotionDialogCN is introduced, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication and achieves an emotion distribution deviation from real human emotion statistics and consistent subject framing, translating into stable unimodal and multimodal performance across acousti...
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning...
Tianchi Liu, Ze-Yang Song, Tian-Rui Wang et al.· 3 citations
This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.
Soma Iwata, K. Inoue, Mu-Yun Wu et al.· Companion Publication of the...· 0 citations
A controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement, and unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration.
Kang-Xiang Xia, Xin-Fa Zhu, Hang-Rui Hu et al.· 0 citations
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...
Yu-Tong Hu, Jin-Ho D. Choi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.