Aug 2026· International Conference on Intelligent Computing· pp. 555-567· 0 citations· 31 references
Computer Science
TL;DR
CETalk is proposed, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control that outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.
Abstract
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.
Speech‐driven 3D facial animation has broad applications in virtual reality, film production, and digital human generation. Personalized style expression plays a crucial role in enhancing realism and expressiveness. However, existing approaches often focus on either emotional expression or speaking style in isolation...
Xemo-Talker is proposed, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision to enhance emotion control.
Chaolong Yang, Yinuo Guo, Kai Yao et al.· 0 citations
Synthesizing holistic co-speech gestures that integrate facial expressions and full-body motion is essential for embodied conversational agents in fields such as virtual reality and film. Existing state-of-the-art approaches predominantly rely on attention-based architectures, which suffer from quadratic computational...
This study presents an integrated multi-modal signal generation framework for intelligent vocalization of traditional folk tales, ensuring emotion-driven speech synthesis with precise lip synchronization. The proposed system models continuous narrative emotion trajectories and embeds them into phoneme-level acoustic ge...
This work presents EMODY Flow, a lightweight flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions.
H. Agarwal, Xavier Alameda-Pineda, Olivier Perrotin· 0 citations
Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving mode...
Kang-Bo Ning, Shan-Shan Gao, Zhao-Qiang Xia et al.· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.