Skip to content
Preprint

MuSeR: Scalable Long-sequence Recommendation with Multi-interest Modeling

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

The contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.

Abstract

Ultra-long user behavior sequences carry rich signals of stable and diverse preferences, yet industrial recommender systems typically truncate histories to a few hundred actions under strict latency and memory budgets, leaving long-term interests under-utilized. Users also pursue multiple heterogeneous intents across modalities such as news, Q&A, and short video, which sparse ID embeddings alone struggle to represent. We present Multi-interest Sequence Representation (MuSeR), a retrieval framework built on the deployed MGS system, which integrates three components: (i) hierarchical temporal compression, which retains recent actions at full resolution while progressively pooling older segments, so that per-user histories of $10^{4}$-$10^{5}$ interactions fit within a fixed serving budget; (ii) disentangled multi-query interest extraction with orthogonality regularization; and (iii) multimodal semantic alignment, which augments sparse item IDs with textual summaries distilled from a large language model. For industrial deployment, MuSeR further adopts asynchronous user-representation refresh with adaptive caching and hierarchical beam-search retrieval across heterogeneous hardware. On three public benchmarks and a large-scale industrial dataset, MuSeR consistently improves Recall@$K$ over strong long-sequence and multi-interest baselines. In online A/B tests on Baidu APP's homepage feed, discovery feed, and short-video scenarios, MuSeR yields +0.26% daily active users and +0.89% total session duration (both statistically significant, p<0.05), alongside reduced serving latency and cost. Rather than proposing a new modeling primitive, our contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.

View source

Similar papers

Book Open access Sep 2026

UniTraj: Cross-Domain Long-Sequence Modeling for Commercial Recommendation

Long-sequence modeling is increasingly important in recommender systems for capturing users’ evolving and long-term interests. In advertising, however, user interaction histories are often highly sparse due to limited exposure opportunities, making ad-only behavior sequences insufficient for effective long-sequence rec...

Xian Hu, Ming Yue, Zhi-Xiang Feng et al. · 0 citations
Book Open access Sep 2026

Interest Sequence for User Modeling in Industrial Short-Form Video Recommendation

Ultra-long user history modeling has been a highly effective approach in modern industrial recommendation systems, with most works heavily utilizing search-based methods and summarization based methods to map massive user interaction logs into latent user representation with algorithm designs to address the scaling cha...

Yuan-Zheng Lin, Diego Uribe Mora, Yuan Shao et al. · 0 citations
#machine learning Preprint Sep 2026

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production...

Jia-Hao Hui, Lin Zhu, Yi-Sheng Hu et al. · 0 citations
Book Open access Sep 2026

Information-Aware Long Sequence Compression for Sequential Recommendation

Sequential recommendation (SR) aims to predict a user’s next interaction by modeling temporal dependencies in historical behavior sequences. However, modeling long sequences introduces two challenges: longer histories often include noisy interactions irrelevant to a user’s core interests, and increasing sequence length...

Woo-Seung Kang, Minje Kim, Suwon Lee et al. · 0 citations
Preprint Sep 2026

ChronicleRec: Pre-training Temporally Anchored Tokens for Lifelong User Modeling

ChronicleRec is a pre-train-and-transfer framework that compresses an ultra-long behavior sequence once into a chronologically ordered set of Chronicle Tokens, which can be cached per user, decoupling ultra-long sequence modeling from online candidate scoring.

Chengkai Huang, Yu-Bin Sheng, Liang Guo et al. · 0 citations
#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation With Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.