HiFTalker: Emotion and Speaking Style Co‐Controllable 3D Facial Animation Via Hierarchical Fusion
Abstract
Speech‐driven 3D facial animation has broad applications in virtual reality, film production, and digital human generation. Personalized style expression plays a crucial role in enhancing realism and expressiveness. However, existing approaches often focus on either emotional expression or speaking style in isolation, overlooking their complementary effects on the naturalness and expressivity of facial animation. To address this limitation, we propose HiFTalker, a novel 3D facial animation framework achieving joint control of emotion and speaking style through hierarchical fusion. We design a visual speaking style modelling module (VSSM) to capture fine‐grained personalized speaking style from input reference video clips and jointly model them with emotional style. Additionally, recognizing that emotional style exerts a global influence on facial expressions while speaking style focuses on fine‐grained local adjustments, we devise a hierarchical fusion method to effectively integrate these two styles. Experiments demonstrate that HiFTalker outperforms existing methods in emotional expression, speaking style modelling, and animation expressiveness.