Skip to content
Preprint

D3ER: Supporting Multi-Modal Recommendation via Disentangle and Distillation-based Dynamic Ensemble

Aug 2026 · 0 citations
Computer Science

TL;DR

A novel method, dubbed Disentangle and Distillation-based Dynamic Ensemble for multi-modal Recommendation (D3ER), which introduces gradient boosting into MR for the first time to formalize the optimization objective for alternately learning HOI and HEI.

Abstract

Incorporating items'information shared among multiple modalities into a fused representation, multi-modal recommendation (MR) has demonstrated documented success than canonical unimodal recommendation. Although several attempts have been made to extract the discriminative information unique in each modality, existing methods suffer from a core limitation: the joint learning of modal-homogeneity discriminative information (HOI) and modal-heterogeneity discriminative information (HEI) tends to weaken their individual effectiveness. To remedy this deficiency, we propose a novel method, dubbed Disentangle and Distillation-based Dynamic Ensemble for multi-modal Recommendation (D3ER). We introduce gradient boosting into MR for the first time to formalize the optimization objective for alternately learning HOI and HEI. This design enables models dedicated to each type of information to focus on their proficient samples, thereby promoting specialized optimization. Furthermore, to mitigate the inherent high storage cost and risk of local optima in gradient boosting, we enhance our framework with knowledge distillation and a global correction regularization. Experiments on prevalent real-world datasets confirm the superiority of our proposed method on MR.

View source

Similar papers

Preprint Sep 2026

LSF-SR: Latent Semantic Fusion for Sequential Recommendation via Flow-based Conditional Variational Autoencoders

Sequential recommendation aims to predict users'future interests from their historical interactions. Although Large Language Models (LLMs) capture rich item semantics, existing methods often struggle to align collaborative signals with textual semantic knowledge. As a result, the learned item representations fail to ca...

Shih-Hong Chen, J. Ying, Vincent S. Tseng · 0 citations
Conference Open access Sep 2026

Three Minds, One Student: Online Multi-Teacher Knowledge Distillation for Multimodal Recommenders

Existing multimodal recommendation models using complex fusion mechanisms (e.g., attention) or multi-stage processes (e.g., early or late fusion) integrate different modalities. However, attention-based adaptive fusion is prone to shortcut learning, where dominant collaborative signals (ID) can overshadow other modalit...

Hang-Tong Xu, Yuanbo Xu, En Wang · 0 citations
Open access Sep 2026

SIHG-Rec: Unleashing the Power of Semantic and Interactive Homogeneous Graphs via Dual-Stage Fusion for Multimodal Recommendation

Recent studies in multimodal recommendation, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered significant interest. Two critical processes in this domain are modality fusion and representation learning. In representation learning, existing studies ofte...

Jin-Feng Xu, Zhe-Yu Chen, Wei Wang et al. · 0 citations
Open access Sep 2026

Cross-Modally Aligned and Temporally Gated Mixture of Experts for Multimodal Sequential Recommendation

Multimodal Sequential recommendation alleviates the semantic insufficiency and data sparsity of item-ID-based models by incorporating side information such as text and images. However, multimodal systems face the dual challenges of feature-space heterogeneity and modality-specific noise, in addition to the dynamic evol...

Yu-Yin Meng, Ai-Xiang Cui, Jun-Lin Zhou et al. · 0 citations
Book Open access Aug 2026

Retrv-MoE: Scaling Unified Multimodal Retrieval with Sparse Mixture-of-Experts

This work proposes Retrv-MoE, a unified retrieval architecture built upon sparse Mixture-of-Experts (MoE), and theoretically and empirically demonstrates that this conditional computation mechanism provides a structural remedy to optimization interference by decoupling the learning trajectories of conflicting tasks and...

Tongxu Lin, Jiayin Xiao · 0 citations
Preprint Aug 2026

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them fo...

Zirui Cheng, Xun Xu, Tiankai Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.