Skip to content

CALSI: Context-Aware Layer Skipping Inference for On-Device LLM Serving

Oct 2026 · IEEE Transactions on Mobile Computing · Vol 25, pp. 17863-17878 · 0 citations · 45 references

Abstract

The large language model (LLM) based on the Transformer architecture and its derived various applications have greatly changed people’s lives. Considering some concerns such as privacy and network conditions, deploying LLM on smart devices has gradually become a research focus. In order to reduce the huge computation and storage overhead of LLMs, many works have studied model compression technology to reduce the model computation and parameter amount, thereby reducing the inference latency. This paper analyzes the characteristics of the on-device LLM service, including small batch size and latency focus, etc. We find the inefficiency of existing model compression technologies and new optimization opportunities, i.e., allocating different layers for different input tokens based on the task QoS requirements. Then we propose CALSI, a context-aware layer skipping LLM inference system for on-device serving. In the offline profiling phase, we analyze the importance of different layers of the model to different tokens and train a lightweight gated predictor. Then, we map the latency QoS requirements of different tasks with the predictor threshold. During the online inference phase, we adaptively allocate the layers that need to be computed to the specific token and design a KV cache delayed computation management mechanism to solve the KV cache missing problem caused by layer skipping. Experiments on real devices show that CALSI can achieve up to 24.1% latency reduction. CALSI has good generalization ability on various models and datasets and is compatible with other existing complementary optimization techniques.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while als...

Qi-Hu Xie, Zi-Wei Li, Yi Kang · 0 citations
#small language model Book Open access Sep 2026

CARE-MoE: Correlation-Aware Expert Placement and Semantic Equivalence Routing for MoE LLM Inference on Edge Devices

CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.

Zhen-Yu Wang, Wei Li, Ao Ren et al. · 0 citations
Preprint Sep 2026

MoSE: Mode-Switching Expander for Mixed LLM Training and Inference

AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static spa...

Fan Yang, Ying Zhou, Bing-Lei Wang et al. · 0 citations
Open access Aug 2026

LLM-Driven Context-Aware Health Monitoring for Resource-Constrained Edge Devices

An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.

Ioannis Tzitzios, A. Dimara, Georgiana Petridou et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.