By providing recommendation services while preserving user data privacy, federated sequential recommendation (FSR) achieves great attention recently. However, FSR suffers from the severe challenge of data sparsity, where item interaction sequences of each user are typically short. The common practice to address the pro...
Yijing Shan, Haozhao Wang, Yi-Chen Li et al.· Proceedings of the 32nd ACM...· 0 citations
A communication-aware adaptive-depth framework is proposed in this paper, termed TrimMoE, which couples layer skipping and confidence-based early exit with substitute execution and server-expert selection under a unified quality budget and proves that the substitution-and-skipping proxy degradation never exceeds the co...
Ning Li, Shuting Bai, Xin Yuan et al.· 0 citations
A similarity-aware expert allocation and distributed deployment framework, dubbed OrderMoE, which aims to accelerate edge MoE inference while balancing inference latency, communication overhead, server workload, and inference quality.
Xin Yuan, Ning Li, Quan Chen et al.· arXiv.org· 2 citations
HetRoute introduces a unified per-assignment cost model that explicitly captures four cost components: cross-server transmission, GPU-CPU offloading, GPU computation with queueing, and quantization-induced quality penalty and establishes fallback feasibility, a bound on the number of participating servers, per-layer op...
Xin Yuan, Ning Li, Wenchao Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.