HAAS: Holistic Attention-free Animation from Speech using Mamba
Abstract
Synthesizing holistic co-speech gestures that integrate facial expressions and full-body motion is essential for embodied conversational agents in fields such as virtual reality and film. Existing state-of-the-art approaches predominantly rely on attention-based architectures, which suffer from quadratic computational complexity and often struggle to model the heterogeneous temporal dynamics of different body parts, leading to motion artifacts and high-frequency jitter. To address these limitations, we propose a novel two-stage, attention-free framework based on Selective State-Space Models (SSMs) for high-fidelity holistic gesture generation. Our approach introduces three technical innovations. First, we propose a hierarchical residual latent mapping using a part-wise Residual-Quantized Variational Motion Prior (RQ-VAE), which decomposes motion into anatomically distinct streams and models them through progressive refinement. Second, we replace the attention bottleneck with a U-Mamba architecture that captures multi-scale temporal dependencies and cross-part interactions with linear computational complexity. Third, the second stage of our model performs continuous residual regression over hierarchical latents to facilitate smooth motion transitions while a learned gating mechanism adaptively fuses speech and text features for part-specific multimodal coherence. Qualitative and quantitative evaluations show that our framework performs highly competitively against baselines, producing more stable and coherent motion with reduced artefacts and jitter. Additionally, our model achieves faster convergence and real-time inference due to the absence of quadratic attention mechanisms. Our results demonstrate that hierarchical residual modelling combined with Mamba-based sequence modelling offers a more efficient and robust solution for holistic human gesture synthesis.