A Hybrid Deep Learning Approach for Emotion Recognition Using Gait Patterns, Body Gestures, and Facial Expressions
Abstract
Human emotion recognition from visual cues has gained significant attention due to its applications in human-computer interaction, surveillance, and healthcare. However, relying on a single modality often limits the robustness of emotion understanding in real-world scenarios. To address this, we propose a novel hybrid deep learning framework that integrates facial expressions, gait patterns, and body gestures for comprehensive emotion recognition. The proposed approach combines convolutional neural networks (CNNs) for facial feature extraction, graph convolutional networks (GCNs) for modeling skeletal and gait dynamics, and spatio-temporal graph convolutional networks (ST-GCNs) for capturing expressive body gestures. To effectively integrate these heterogeneous modalities, a transformer-based cross-attention mechanism is employed, enabling adaptive feature fusion by learning inter-modal relationships. Furthermore, temporal dependencies across sequences are modeled using self-attention, allowing the network to capture both short- and long-range emotional dynamics. Extensive experiments demonstrate that the proposed CNN + GCN + Transformer hybrid model significantly improves recognition performance compared to unimodal and conventional multimodal approaches. The results highlight the effectiveness of combining spatial, structural, and temporal representations through attention-driven fusion for robust and accurate emotion recognition.