Skip to content

LLM Deployment Strategies on Mobile Edge Servers for Dynamic Uncertain User Requests

2026 · IEEE Transactions on Network and Service Management · Vol 23, pp. 5832-5850 · 0 citations · 43 references
Computer Science

Abstract

Leveraging on the task planning and solving capability of pretrained Large Language Models (LLMs), deploying LLM agents on Mobile Edge Computing (MEC) edge servers brings significant benefits for an Internet of Things (IoT) network for providing enhanced AI intelligence with acceptable delay. In this work, we consider the edge LLMs deployment strategy in an end-edge-cloud LLM agents system for the IoT services, which jointly determines the locations and number of LLM initializations and user requests offloading strategy in a dynamic network environment with stochastic user requests. We formulate this joint LLM Deployment and inference Tasks Offloading (LLMDTO) problem. Typically, we design an LLM service performance evaluation mechanism by measuring its processing delay with stochastic user requests arrivals by Stochastic Network Calculus (SNC). Due to the complexity of the LLMDTO problem, we decompose this joint optimization problem into two subproblems and propose an algorithm based on Multi Agent Deep Reinforcement Learning (MADRL) scheme. To accelerate the training process of the DRL, a reward model is designed by applying the Kolmogorov Arnold Networks (KAN) to return a fast reward estimation. Finally, we validate the proposed algorithm through extensive simulations and results show the effectiveness of the proposition on lower deployment cost and delay in a dynamic network environment.

View source

Similar papers

Open access 2026

Cooperative Task Offloading in Mobile Edge Computing via an Improved MASAC Framework

An adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC, which uses a Beta policy to model bounded actions and achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings.

Zheng Yao, Jie Liu, Changjun Deng et al. · 0 citations
Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations
Open access Jul 2026

Task-Offloading Optimization in Mobile Edge Computing for Smart Library Services

A preference-adaptive dueling double deep Q-network algorithm, termed PA-DDQN, is proposed by integrating preference conditioning, multi-head attention, a dueling architecture, and double Q-learning, demonstrating its effectiveness in enhancing service responsiveness, energy efficiency, and reliability in smart library MEC systems.

Jingjing Qu, Peiying Zhang, Ruixin Wang et al. · 0 citations
Aug 2026

DNN task computation offloading and resource allocation optimization strategy based on probabilistic early exit

A Mixed Integer Nonlinear Programming (MINLP) model with the objective of a weighted sum of long-term average task completion rate, total latency and energy consumption is established, which improves the task completion rate by 4% in high load scenarios and achieves a better balance between latency and energy consumption.

Xianzhong Tian, Xuhua Mao, Xipeng Zhou · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations
Jul 2026

Intelligent Placement of 5G Network Functions on Edge-Based Infrastructures

A constrained optimization model that supports different management goals through alternative objective functions (latency-aware or power-aware) while enforcing operational constraints, including node capacities, slice-specific latency bounds, and explicit limits on VNF migrations/relocations between scheduling periods is proposed.

R. Moreno-Vozmediano, E. Huedo, R. Montero et al. · 0 citations