Multi-Scale Spatiotemporal EEG and Self-Supervised Audio Fusion: A Mixture-of-Experts Approach to Continuous Affect
Abstract
Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal EEG encoder that captures per-channel temporal dynamics and inter-channel dependencies and (ii) self-supervised audio embeddings extracted from SSLAM in a feature-based transfer setting. To integrate modalities robustly, we introduce a modality-sensitive mixture-of-experts (MoE) fusion module with modality-specific experts and a learned gating network that dynamically weights EEG and audio contributions. Under session-preserving partitioning with 5-fold cross-validation, the proposed approach achieves strong agreement with ground-truth valence trajectories (PCC=0.778, CCC=0.742, RMSE=0.034), comparing favorably with representative prior results reported under broadly related MAHNOB-HCI protocols. Finally, we provide interpretability evidence by linking the EEG encoder’s alpha-band saliency asymmetry to classical alpha-band power asymmetry patterns, thereby supporting the neurophysiological plausibility of the learned representations.