Cross-Modally Aligned and Temporally Gated Mixture of Experts for Multimodal Sequential Recommendation
Abstract
Multimodal Sequential recommendation alleviates the semantic insufficiency and data sparsity of item-ID-based models by incorporating side information such as text and images. However, multimodal systems face the dual challenges of feature-space heterogeneity and modality-specific noise, in addition to the dynamic evolution of user interests over time. Existing methods still struggle to jointly handle cross-modal alignment and time-aware preference modeling. To address these challenges, we propose a multimodal sequential recommendation framework with cross-modal alignment and temporal gating, which leverages item ID, text, and image modalities to capture users’ dynamic interests. The proposed model contains three core components. First, a cross-modal alignment mixture-of-experts module preserves modality-specific features with dedicated experts and captures shared semantics with common experts, thereby mitigating the semantic mismatch inherent in direct fusion. Second, a hierarchical time-aware mixture-of-experts module uses short-term intervals, long-term spans, and periodic time encodings for expert routing, and applies a time-aware modality gate to adaptively adjust the importance of ID, text, and image modalities under different temporal contexts. Third, a sequential interest contrastive learning objective enhances the discriminability of ID-based sequential interest representations by leveraging dynamic temperature scaling, multi-scale positive samples, hard negative mining, and diversity regularization. Experiments on games, beauty, and toys demonstrate that the proposed method consistently outperforms representative sequential and multimodal recommendation baselines on Normalized Discounted Cumulative Gain (NDCG)@5, NDCG@10, Mean Reciprocal Rank (MRR)@5, and MRR@10. Furthermore, ablation results validate the effectiveness of each proposed component.