Skip to content
Open access

Lightweight and robust audio-visual emotion recognition via multi-scale mamba temporal modeling and quality-aware expert fusion

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 48 references

Abstract

Multimodal emotion recognition for real-time human–computer interaction requires both high recognition accuracy and low computational cost. However, existing audio-visual methods often rely on heavy Transformer-based fusion or high-capacity visual backbones, making them difficult to deploy on edge devices. Moreover, their robustness to degraded audio-visual inputs and their class-wise behavior for ambiguous emotions remain insufficiently analyzed. To address these issues, we propose a lightweight audio-visual emotion recognition framework. The visual stream uses ShuffleNet for efficient facial feature extraction, while the audio stream uses multiscale MFCCs and efficient sequence modeling to capture emotional prosody. A cross-modal auxiliary fusion module is further introduced to align audio and visual representations, and lightweight channel attention is used to emphasize emotion-relevant features. Experiments on CREMA-D and IEMO-CAP demonstrate that the proposed method achieves competitive recognition accuracy with significantly fewer parameters and lower computational cost. Additional analyses on edge-device inference, audio-visual degradation, class-wise confusion, and batch-size sensitivity further validate the efficiency and robustness of the proposed framework.

Read PDF