SPLGMMamba: A Self-Paced Learning-Guided Multimodal Mamba Framework for Depression Recognition
Abstract
Depression manifests through heterogeneous behavioral cues, and multimodal methods can improve its recognition by integrating complementary information across modalities. However, existing approaches remain limited by the high computational cost of long-sequence modeling, insufficient cross-modal interaction, and interference from redundant features. Variations in modality contributions across samples may also cause modality dominance and unstable optimization. To address these issues, this paper proposes a Self-Paced Learning Guided Multimodal Mamba framework, termed SPLGMMamba, for audio-visual depression severity prediction. SPLGMMamba extracts modality-specific representations using a Video Spatio-Temporal Mamba Network (VSTMNet) and an Audio Temporal Mamba Network (ATMNet). A Multimodal Fusion Network (MFNet) performs efficient cross-modal interaction in a unified latent state space with linear computational complexity. A Self-Paced Learning Network (SPLNet) further regulates contribution-aware sample inclusion and modality modulation according to dynamically estimated modality-contribution patterns. Experimental results on the AVEC 2013, AVEC 2014, and AVEC 2017 datasets show that SPLGMMamba achieves the best multimodal performance, with MAE/RMSE values of 6.06/7.03, 5.10/6.61, and 4.55/5.38, respectively. These results indicate that combining state-space-based multimodal fusion with contribution-aware self-paced optimization provides an effective approach to audio-visual depression severity prediction.