Skip to content

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

Jul 2026 · arXiv.org · Vol abs/2607.15755 · 1 citation · 57 references
Computer Science

TL;DR

This work proposes AuEmoChat, a CSS framework for authentic emotion understanding and rendering, and develops AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories.

Abstract

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. Furthermore, we propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech. The code and speech demos will be available at: https://github.com/AI-S2-Lab/AuEmoChat.

View source

Similar papers

Preprint Aug 2026

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

EmotionDialogCN is introduced, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication and achieves an emotion distribution deviation from real human emotion statistics and consistent subject framing, translating into stable unimodal and multimodal performance across acousti...

Yi Zheng, Yifan Xu, Yan Zhou et al. · 0 citations
Preprint Aug 2026

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning...

Tianchi Liu, Ze-Yang Song, Tian-Rui Wang et al. · 3 citations
Book Open access Sep 2026

Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues

This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.

Soma Iwata, K. Inoue, Mu-Yun Wu et al. · 0 citations
#natural language process... Preprint Sep 2026

From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS

A controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement, and unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration.

Kang-Xiang Xia, Xin-Fa Zhu, Hang-Rui Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue

Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...

Yu-Tong Hu, Jin-Ho D. Choi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.