Serving offline large language model (LLM) inference workloads (e.g., log summarization and bulk translation) can consume up to 30% of GPUs in production. Despite this significant share, the characteristics of offline inference remain largely understudied. In this paper, we start by analyzing 1.5 million tasks comprisi...
Le-Ping Yang, Xue Li, Kun Qian et al.· Proceedings of the ACM SIGOP...· 0 citations
Recommender systems deployed at scale are predominantly organized around platform-centric architectures in which each service independently constructs and optimizes an internal representation of the user. As large language models (LLMs) become upstream entry points to digital services, personalization may be reorganize...
Jia-Hao Liu, Ming-Zhe Han, Guan Liu et al.· Proceedings of the 20th ACM...· 0 citations
Federated recommendation enables collaborative model training while keeping user interaction data on local clients. A central problem in federated recommendation is how to aggregate useful information across clients for personalized recommendation. Existing personalized aggregation methods usually construct client rela...
Ming-Zhe Han, Jia-Hao Liu, Dong-Sheng Li et al.· 0 citations
This work surveys dozens of recent works that report compression results on real hardware and extracts practical deployment guidelines from them, and deploys compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation.
Subhransu Das, Jiaming Cheng, Arnav Kumar et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.