The contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.
Abstract
Ultra-long user behavior sequences carry rich signals of stable and diverse preferences, yet industrial recommender systems typically truncate histories to a few hundred actions under strict latency and memory budgets, leaving long-term interests under-utilized. Users also pursue multiple heterogeneous intents across modalities such as news, Q&A, and short video, which sparse ID embeddings alone struggle to represent. We present Multi-interest Sequence Representation (MuSeR), a retrieval framework built on the deployed MGS system, which integrates three components: (i) hierarchical temporal compression, which retains recent actions at full resolution while progressively pooling older segments, so that per-user histories of $10^{4}$-$10^{5}$ interactions fit within a fixed serving budget; (ii) disentangled multi-query interest extraction with orthogonality regularization; and (iii) multimodal semantic alignment, which augments sparse item IDs with textual summaries distilled from a large language model. For industrial deployment, MuSeR further adopts asynchronous user-representation refresh with adaptive caching and hierarchical beam-search retrieval across heterogeneous hardware. On three public benchmarks and a large-scale industrial dataset, MuSeR consistently improves Recall@$K$ over strong long-sequence and multi-interest baselines. In online A/B tests on Baidu APP's homepage feed, discovery feed, and short-video scenarios, MuSeR yields +0.26% daily active users and +0.89% total session duration (both statistically significant, p<0.05), alongside reduced serving latency and cost. Rather than proposing a new modeling primitive, our contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.
Long-sequence modeling is increasingly important in recommender systems for capturing users’ evolving and long-term interests. In advertising, however, user interaction histories are often highly sparse due to limited exposure opportunities, making ad-only behavior sequences insufficient for effective long-sequence rec...
Xian Hu, Ming Yue, Zhi-Xiang Feng et al.· Proceedings of the 20th ACM...· 0 citations
Ultra-long user history modeling has been a highly effective approach in modern industrial recommendation systems, with most works heavily utilizing search-based methods and summarization based methods to map massive user interaction logs into latent user representation with algorithm designs to address the scaling cha...
Yuan-Zheng Lin, Diego Uribe Mora, Yuan Shao et al.· Proceedings of the 20th ACM...· 0 citations
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production...
Jia-Hao Hui, Lin Zhu, Yi-Sheng Hu et al.· 0 citations
Sequential recommendation (SR) aims to predict a user’s next interaction by modeling temporal dependencies in historical behavior sequences. However, modeling long sequences introduces two challenges: longer histories often include noisy interactions irrelevant to a user’s core interests, and increasing sequence length...
Woo-Seung Kang, Minje Kim, Suwon Lee et al.· Proceedings of the 20th ACM...· 0 citations
ChronicleRec is a pre-train-and-transfer framework that compresses an ultra-long behavior sequence once into a chronologically ordered set of Chronicle Tokens, which can be cached per user, decoupling ultra-long sequence modeling from online candidate scoring.
Chengkai Huang, Yu-Bin Sheng, Liang Guo et al.· 0 citations
PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.
Hyunsoo Na, Minseok Gang, Sang-goo Lee et al.· ACM Transactions on Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.