Skip to content
Book Open access

Multi-Scale Spatiotemporal EEG and Self-Supervised Audio Fusion: A Mixture-of-Experts Approach to Continuous Affect

Oct 2026 · 0 citations

Abstract

Emotion-aware human–computer interaction increasingly relies on continuous emotion recognition (CER) to track affective states over time. This paper investigates continuous valence prediction on the MAHNOB-HCI database using a multimodal EEG+audio framework. The proposed model combines (i) a hierarchical spatiotemporal EEG encoder that captures per-channel temporal dynamics and inter-channel dependencies and (ii) self-supervised audio embeddings extracted from SSLAM in a feature-based transfer setting. To integrate modalities robustly, we introduce a modality-sensitive mixture-of-experts (MoE) fusion module with modality-specific experts and a learned gating network that dynamically weights EEG and audio contributions. Under session-preserving partitioning with 5-fold cross-validation, the proposed approach achieves strong agreement with ground-truth valence trajectories (PCC=0.778, CCC=0.742, RMSE=0.034), comparing favorably with representative prior results reported under broadly related MAHNOB-HCI protocols. Finally, we provide interpretability evidence by linking the EEG encoder’s alpha-band saliency asymmetry to classical alpha-band power asymmetry patterns, thereby supporting the neurophysiological plausibility of the learned representations.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.