Skip to content
Preprint

Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

Xemo-Talker is proposed, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision to enhance emotion control.

Abstract

Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion-related losses across the entire motion space poses significant difficulties due to the inherent trade-off between accurate lip synchronization and fine-grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less-principal components achieves a better emotion-lip synchronization balance, as principal components mainly encode high-energy articulation and pose variations. Building on this insight, we propose Xemo-Talker, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision. To enhance emotion control, we design a Tri-Loss consisting of inter-class separation, intra-class compactness, and less-principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo-Talker.

View source

Similar papers

Open access Aug 2026

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

CETalk is proposed, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control that outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

Peng Jia, Li Dai, Zhen Xiao et al. · 0 citations
Open access Aug 2026

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis.

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling space for facial features, ensuring that they a) operate independently without...

Shuai Tan, Bin Ji, Ye Pan · 0 citations
Preprint Aug 2026

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning...

Tianchi Liu, Ze-Yang Song, Tian-Rui Wang et al. · 3 citations
#machine learning Preprint Aug 2026

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

This work presents EMODY Flow, a lightweight flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions.

H. Agarwal, Xavier Alameda-Pineda, Olivier Perrotin · 0 citations
Preprint Aug 2026

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional...

Yi Zheng, Yifan Xu, Yan Zhou et al. · 0 citations
Open access Sep 2026

Zero-shot emotional speech synthesis based on feature decoupling and adaptive loss-threshold reweighting

Deep learning-based zero-shot speech synthesis has achieved substantial progress in speaker generalization, but stable modeling remains challenging in fine-grained emotional scenarios. Existing systems often process textual and emotional conditions through shared or closely coupled pathways, which may introduce interfe...

Yan Zhu, Yu Wang, Yi-Jin Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.