Skip to content
Open access

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

Aug 2026 · International Conference on Intelligent Computing · pp. 555-567 · 0 citations · 31 references
Computer Science

TL;DR

CETalk is proposed, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control that outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

Abstract

Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

Read PDF

Similar papers

Sep 2026

HiFTalker: Emotion and Speaking Style Co‐Controllable 3D Facial Animation Via Hierarchical Fusion

Speech‐driven 3D facial animation has broad applications in virtual reality, film production, and digital human generation. Personalized style expression plays a crucial role in enhancing realism and expressiveness. However, existing approaches often focus on either emotional expression or speaking style in isolation...

Xu Yang, Shi-Guang Liu, Qing Xu · 0 citations
Preprint Aug 2026

Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

Xemo-Talker is proposed, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision to enhance emotion control.

Chaolong Yang, Yinuo Guo, Kai Yao et al. · 0 citations
Book Open access Oct 2026

HAAS: Holistic Attention-free Animation from Speech using Mamba

Synthesizing holistic co-speech gestures that integrate facial expressions and full-body motion is essential for embodied conversational agents in fields such as virtual reality and film. Existing state-of-the-art approaches predominantly rely on attention-based architectures, which suffer from quadratic computational...

Laxmi Narayen Nagarajan Venkatesan, Harsh Vardhan Singh, Vansh Sinha et al. · 0 citations
Open access Aug 2026

Intelligent Interpretation of Traditional Folk Tales through Vocalization: a Generative Model Based on Emotion-Driven and Lip-Syncing

This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic ge...

S. Li · 0 citations
#machine learning Preprint Aug 2026

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

This work presents EMODY Flow, a lightweight flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions.

H. Agarwal, Xavier Alameda-Pineda, Olivier Perrotin · 0 citations
Open access Aug 2026

Parameter-Efficient Audio-Visual Dynamic Facial Expression Recognition with Mamba Fusion Adapters and Frame-Level Feature Arrangement

Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving mode...

Kang-Bo Ning, Shan-Shan Gao, Zhao-Qiang Xia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.