Skip to content

Author

Xiaoliang Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Lightweight and robust audio-visual emotion recognition via multi-scale mamba temporal modeling and quality-aware expert fusion

Multimodal emotion recognition for real-time human–computer interaction requires both high recognition accuracy and low computational cost. However, existing audio-visual methods often rely on heavy Transformer-based fusion or high-capacity visual backbones, making them difficult to deploy on edge devices. Moreover, their robustness to degraded audio-visual inputs and their class-wise behavior for ambiguous emotions remain insufficiently analyzed. To address these issues, we propose a lightweight audio-visual emotion recognition framework. The visual stream uses ShuffleNet for efficient facial feature extraction, while the audio stream uses multiscale MFCCs and efficient sequence modeling to capture emotional prosody. A cross-modal auxiliary fusion module is further introduced to align audio and visual representations, and lightweight channel attention is used to emphasize emotion-relevant features. Experiments on CREMA-D and IEMO-CAP demonstrate that the proposed method achieves competitive recognition accuracy with significantly fewer parameters and lower computational cost. Additional analyses on edge-device inference, audio-visual degradation, class-wise confusion, and batch-size sensitivity further validate the efficiency and robustness of the proposed framework.

Tianxing Zhang, Hadi Affendy Bin Dahlan, Fadhilah Rosdi et al. · 0 citations