Oct 2026· IEEE Transactions on Mobile Computing· Vol 25, pp. 17863-17878· 0 citations· 45 references
Abstract
The large language model (LLM) based on the Transformer architecture and its derived various applications have greatly changed people’s lives. Considering some concerns such as privacy and network conditions, deploying LLM on smart devices has gradually become a research focus. In order to reduce the huge computation and storage overhead of LLMs, many works have studied model compression technology to reduce the model computation and parameter amount, thereby reducing the inference latency. This paper analyzes the characteristics of the on-device LLM service, including small batch size and latency focus, etc. We find the inefficiency of existing model compression technologies and new optimization opportunities, i.e., allocating different layers for different input tokens based on the task QoS requirements. Then we propose CALSI, a context-aware layer skipping LLM inference system for on-device serving. In the offline profiling phase, we analyze the importance of different layers of the model to different tokens and train a lightweight gated predictor. Then, we map the latency QoS requirements of different tasks with the predictor threshold. During the online inference phase, we adaptively allocate the layers that need to be computed to the specific token and design a KV cache delayed computation management mechanism to solve the KV cache missing problem caused by layer skipping. Experiments on real devices show that CALSI can achieve up to 24.1% latency reduction. CALSI has good generalization ability on various models and datasets and is compatible with other existing complementary optimization techniques.
Hydra is presented, a common-schema, phase-aware workload characterization framework for LLM inference on edge SoCs that enables reproducible, phase-aware characterization of edge LLM inference.
Amir Taherin, Sana Taghipour Anvari, Charles Amante et al.· 2 citations
Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while als...
CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.
Zhen-Yu Wang, Wei Li, Ao Ren et al.· Proceedings of the Internati...· 0 citations
AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static spa...
Fan Yang, Ying Zhou, Bing-Lei Wang et al.· 0 citations
This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference, and conducts a broad empirical study covering prompt properties, datasets, models, and devices.
Wei-Hao Hu, Xiaolong Tu, Dawei Chen et al.· 0 citations
An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.
Ioannis Tzitzios, A. Dimara, Georgiana Petridou et al.· Electronics· 0 citations
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.