A transferable residual adapter that injects additional ranking-specific features into ranking in a residual manner and an asymmetric multi-epoch training strategy that resets sparse parameters while continuously accumulating dense parameters across epochs, alleviating the overfitting of sparse parameters are proposed.
Abstract
Transformers have shown promising performance in LLMs due to their outstanding scalability, several studies have investigated the scalability of Transformers for industrial recommendation. They typically rely on a single ranking model to optimize both sparse and dense parameters from scratch, resulting in substantial computational resource consumption and slow convergence. Fortunately, the pre-training models offer an effective solution to the above issues by providing favorable initialization of both sparse and dense parameters for the subsequent ranking. However, they still face two major limitations: (1) Since the input features used in pre-training and ranking are usually inconsistent, directly transferring dense parameters from pre-training to ranking may lead to negative transfer. (2) Multi-epoch training during the ranking process may result in the overfitting of sparse parameters, while freezing the sparse parameters limits their adaptability to the ranking objectives. To this end, we propose a Scaling Transformer for Industrial Recommendation via Transferable Generative Pre-training, termed LazFormer. Specifically, we first present a generative pre-training module to autoregressively generate sequential features, providing favorable initialization of both sparse and dense parameters for the subsequent ranking. To solve the negative transfer of dense parameters, we propose a transferable residual adapter that injects additional ranking-specific features into ranking in a residual manner. Moreover, a request-aware ranking module integrates long-sequence compression, hybrid sparse attention, and a request-aware paradigm to efficiently model users'long sequences. Besides, we further propose an asymmetric multi-epoch training strategy that resets sparse parameters while continuously accumulating dense parameters across epochs, alleviating the overfitting of sparse parameters.
ReST is proposed, a recommendation-native Transformer scaling framework that achieves higher accuracy and scales more consistently along sequence length, depth, and width, where LLM-style Transformer blocks saturate.
Jie Chen, Xiang-Qia Yu, Yan-Chao Lian et al.· 0 citations
REP-LIE leverages the gradients of LoRA low-rank matrices to estimate the importance of weights without requiring full gradient computation, and a stability score is introduced, serving as the basis for iterative pruning of unimportant model parameters.
Peng Liu, Hui-Bing Zeng, Yi-Qun Zhang et al.· IEEE Transactions on Emergin...· 0 citations
TransRetrieval is presented, a Transformer-based retrieval framework that scales with both computational budget and cross-domain data and introduces weighted average aggregation, which restores the homogeneous-token assumption Transformers rely on, and target token compression that cuts per-candidate FLOPs while preser...
Zhi-Fei Zheng, Yun-Fei Liu, Bin Liu et al.· 0 citations
Large language models (LLMs) have demonstrated strong capabilities in recommendation tasks such as item, sequence, conversational recommendation, and explanation generation. However, LLM weights are typically shared across all users. Adapting these models to individual users remains a fundamental challenge that require...
Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al.· Proceedings of the 20th ACM...· 0 citations
SplitLite is proposed, a communication-efficient split federated LoRA fine-tuning method that exploits the low effective rank structure of consecutive-epoch activation and gradient residuals, thereby significantly reducing both activation uplink and gradient downlink traffic.
PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.
Hyunsoo Na, Minseok Gang, Sang-goo Lee et al.· ACM Transactions on Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.