Skip to content

Category

edge computing

757 papers

#edge computing Sep 2026

Efficient Layer-Granularity Unloading for LLMs in Edge Computing

Advancements in edge computing and container technology have made it increasingly popular and convenient to deploy Large Language Models (LLMs) through containers at the edge. However, the limited GPU resources of edge servers make it impractical to retain the model in GPU memory for long periods due to the high memory cost, especially when they remain idle without user requests. Existing work unloads the entire idle models to reduce memory costs on edge servers, but reloading them introduces significant loading delays that affect task Quality of Service (QoS). Therefore, efficient management of idle models is a critical issue that has been largely neglected in existing research and requires urgent attention. To address this gap, this paper studies the problem of idle model management from the perspective of the trade-off between memory cost and loading delay under the QoS constraint. A novel layer-granularity model unloading method is proposed, which leverages the layered characteristics of the model. We formulate an online joint optimization problem to determine which layers to unload and when, and present a layer-granularity unloading strategy inspired by the ski rental problem to solve it. We implement a real system with layer-granularity unloading for LLMs on NVIDIA GPUs and validate the effectiveness of the proposed method. Experimental results show it effectively trades off memory cost and loading delay, improving overall performance by up to 39.6%.

Zhenzheng Li, Zhiqing Tang, Jianxiong Guo et al. · 1 citation
#edge computing Sep 2026

Joint Latency and Charge Cost Minimization for Reliable Task Offloading in Dispersed Computing: A Multi-Objective Optimization Approach

Dispersed computing has emerged as a promising paradigm that leverages underutilized resources from massive Internet of Things devices (IoTDs) to enhance the computing capacity at the network edge. However, existing works about the dispersed computing overlook the heterogeneous computing environment with parallel and serial computations and task reliability requirements for the hardware-constrained IoTDs, and they lack multi-objective optimization approaches to optimize the task offloading. To address the challenges, we propose a comprehensive scheme to achieve a delay-aware and economic-aware dispersed computing paradigm by using a multi-objective optimization approach. Particularly, we consider parallel processing at an edge server and serial processing at the lightweight IoTDs, and leverage the task redundancy to satisfy the task reliability requirements on the IoTD side. We further formulate a constrained multi-objective optimization problem (CMOP) aiming at jointly optimizing the task assignment, bandwidth allocation, and CPU frequency allocation to simultaneously minimize the total delay cost and the total charge cost of the tasks. To address the CMOP, we propose an improved constrained multi-objective evolutionary algorithm that employs a dual-population cooperative mechanism between two populations and a repairing constraint-handling technique. The dual-population cooperative mechanism can balance convergence toward Pareto optimality and solution diversity maintenance. The repairing constraint-handling technique is designed to guide solutions toward feasible regions, achieving efficient exploration of complex constrained search spaces. Simulation results demonstrate the superiority of our algorithm in seeking the better-converged and better-distributed Pareto optimal solutions to well address the tradeoffs between the two objectives.

Xumin Huang, Zexiong Wu, Chaoda Peng et al. · 2 citations
#edge computing Sep 2026

Energy-Efficient Task Allocation for Green Aerial Edge Computing Based on Metaverse Users: A Mean Field Game Approach

We consider the energy-constrained task allocation problem in large-scale Aerial Edge Computing (AEC) systems, which encompasses a series of tightly coupled decision-making processes, including which tasks need to be processed by uncrewed aerial vehicles (UAVs), how to allocate these tasks and balance energy across UAVs for delay-sensitive requirements. However, little attention has been devoted to exploring the above coupled decision-making problem in AEC with various resource and energy constraints, which is further complicated by energy dynamics (UAV battery states), task-specific consumption, and allocation-feedback balance. In this paper, we formulate a multi-dimensional joint optimization problem, simultaneously optimizing task allocation and energy rewarding to maximize long-term system rewards while balancing service quality and energy efficiency. To this end, we propose a green aerial edge computing framework where partial UAVs are equipped with energy harvesting modules to collect ambient energy. To circumvent the intractable computational complexity arising from the coupled energy states of massive UAVs, we design a distributed solution method based on the mean field game, which decouples the dense multi-agent interactions into a game between an individual UAV and the aggregate population state, thereby transforming the complex global optimization problem into a set of equivalent scalable subproblems. We develop an optimal energy valuation scheme to guide UAV behavior. Numerical results show that our mechanism can effectively ensure sustainable system operation while maintaining high quality of service for metaverse users, outperforming existing methods in both system sustainability and service responsiveness.

Lianbo Ma, Dingsige Chen, Yuee Zhou et al. · 0 citations
#edge computing Sep 2026

Task completion-oriented service migration for connected autonomous vehicles in multi-server edge computing

To address the urgent practical challenge of service migration for connected autonomous vehicles (CAVs) in mobile edge computing (MEC), this study aims to maximize the task completion rate, particularly for safety-critical operations. Existing approaches often overlook the completion status of tasks with different priority levels during frequent service migrations and fail to co-optimize multiple constraints such as energy consumption, latency, and offloading cost. Consequently, it remains difficult to reliably complete highly urgent tasks in resource-constrained edge environments. To tackle this issue, we propose a comprehensive two-stage solution: the Improved Task Offloading and Service Migration (ITOSM) algorithm. In the first stage, a weighted evaluation model based on information entropy is constructed by integrating transmission time, execution time, and offloading cost. Tasks are offloaded to the edge server with either the highest or second-highest weighted sum according to their urgency level. In the second stage, service migration decisions are optimized using an improved binary particle swarm optimization (BPSO) algorithm with enhanced local search capability. Experimental results demonstrate that ITOSM outperforms existing methods, achieving up to 10.00% higher completion rates for Extremely Important Tasks (EITs) with strict deadlines. This improvement directly contributes to safer and more reliable CAV operations, highlighting the practical significance of this work for intelligent transportation systems.

Jing Liu, Jie-Yi Deng, Longxin Zhang et al. · 0 citations
#edge computing Sep 2026

Truthful Online Double Auction-Based Resource Allocation Mechanisms for Partial Computation Offloading in Collaborative Edge Computing

As mobile applications become increasingly computation-intensive, mobile devices (MDs) face growing limitations due to their constrained computational capabilities and battery life. Collaborative Edge Computing (CEC) has emerged as a promising solution to address these challenges by enabling multiple edge service providers (ESPs) to offer computation offloading services to MDs. As such, a CEC resource trading market is essential for efficient interactions between MDs and ESPs. However, jointly determining the offloading ratios, allocating combinatorial computation and communication resources, and designing appropriate pricing strategies in a dynamic market remains a significant challenge. To this end, we propose a truthful online double auction-based resource allocation mechanism for partial computation offloading (TRAPO) that explicitly accounts for the stochastic nature of both MDs and ESPs. TRAPO first leverages spatial diversity to construct a set of bids for each MD by mapping their task requirements into resource demands through considering MDs’ preferences and partial offloading. Next, we match resource-demanding MDs with resource-supplying ESPs based on adaptive valid price thresholds to maximize social welfare, and calculate the payments of MDs and the rewards of ESPs. Theoretical analyses demonstrate that TRAPO satisfies truthfulness, budget balance, individual rationality, and computational tractability. Simulation experiments further verify the effectiveness and efficiency of TRAPO.

Dongkuo Wu, Xingwei Wang, Xueyi Wang et al. · 0 citations
#edge computing Sep 2026

PreSFC: Predictive SFC Migration via Multi-Slot Mobility Forecasting in MEC Networks

Network Function Virtualization (NFV) is a foundational technology for Mobile Edge Computing (MEC). It delivers network services by chaining Virtual Network Functions (VNFs) into sequential Service Function Chains (SFCs). One of the most critical challenges in MEC is how to provide continuous and stable services to high-mobility user, such as intelligent vehicles and drones. However, current mobility-aware SFC migration methods remain constrained by either post-hoc reaction or myopic prediction horizons, failing to reconcile the divergent timescales of network services and user mobility, thus resulting in suboptimal resource allocation and service delivery. In this paper, we first formulate the predictive mobility-aware SFC migration problem as an NP-hard Integer Linear Programming (ILP) problem. Aiming to mitigate service disruption for mobile users in MEC networks, we propose PreSFC, a predictive SFC migration framework that integrates multi-slot mobility forecasting with fine-grained network state tracking. We first design Gformer, a deep learning-based long sequence time-series forecasting model, which operates on short time slots (less than 200 ms) to sensitively capture network dynamics while predicting over multiple slots (e.g., 50 slots) to effectively track user mobility. This dual-scale design explicitly addresses the temporal disparity between mobility patterns and service requirements. Based on the predictions, we further propose an Optimal Sub-period Partitioning Migration (OSPM) algorithm to determine migration timing and locations. Extensive simulations show that our approach reduces the maximum and average downtime by approximately 55% and 40%, respectively, compared to benchmark methods.

Ji Li, Songtao Guo, Quanjun Zhao et al. · 0 citations
#edge computing Sep 2026

Service Enhancement and Reliability Assurance in 6G Vehicular Networks via a Stackelberg Game-Theoretic Approach

With the rapid development of 6G and Internet of Vehicles (IoV) technologies, the volume of computation-intensive tasks generated by intelligent vehicles is growing exponentially. Given limited onboard processing capabilities, vehicles increasingly rely on edge servers deployed by service providers (SPs) at roadside units to offload tasks. Vehicle clients can offload the tasks to SPs to mitigate their onboard computation load, while SPs derive economic benefits through the provision of computation resources. However, this interaction introduces a conflict of interest, as vehicles aim to minimize their offloading costs, while SPs seek to maximize revenue. To address this problem, we propose SPOR, a Stackelberg game-based service priority-aware computation offloading and resource pricing scheme in IoV. SPOR is a hierarchical game-theoretic framework in which SPs act as leaders setting prices, while vehicles act as followers determining their offloading strategies. A novel service prioritization function is introduced, incorporating booking price, system load, and reputation to ensure fair and balanced resource allocation. We provide a theoretical proof of the existence and uniqueness of a Nash equilibrium. Extensive experiments on a real-world vehicle edge computing dataset show that SPOR outperforms baseline methods in delay, energy consumption, average load, and task completion rate. Notably, SPOR maintains task completion rates above 97% even under heavy workloads, demonstrating its effectiveness in enhancing system reliability and overall performance.

Kai Peng, Yuanlin Lin, Shuai Zhao et al. · 2 citations
#edge computing Sep 2026

MERA: A Green Edge Resource Control System With Privacy-Preservation via Mean-Field Reinforcement Learning

The global rollout of 5G networks has spurred the rapid deployments of edge servers for hosting latency-sensitive web applications, which improves quality of experience (QoE). However, current efforts fall short in the substantial energy costs associated with the 24/7 operation of edge servers and overlook user privacy by requiring accurate user information for service provision, eroding the sustainability of multi-access edge computing (MEC). To enhance the QoE and service performance while ensuring privacy in MEC, we systematically formulate the interaction among edge servers as a privacy-preserving experience-aware edge resource control (PEERC) problem. To address this, we conduct a global resource control and propose a collaborative resource allocation system named MERA. MERA leverages <inline-formula><tex-math notation="LaTeX">$k$</tex-math><alternatives><mml:math><mml:mi>k</mml:mi></mml:math><inline-graphic xlink:href="xia-ieq1-3705464.gif"/></alternatives></inline-formula>-anonymity data obfuscation to protect user location and resource demand privacy while enhancing service performance and energy efficiency with mean-field multi-agent reinforcement learning. Extensive experiments based on a synthetic real-world dataset demonstrate that MERA significantly surpasses benchmarks in terms of QoE, user coverage, privacy, and energy efficiency by <inline-formula><tex-math notation="LaTeX">$1.18\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>18</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq2-3705464.gif"/></alternatives></inline-formula>, <inline-formula><tex-math notation="LaTeX">$1.24\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>24</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq3-3705464.gif"/></alternatives></inline-formula>, <inline-formula><tex-math notation="LaTeX">$1.63\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>63</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq4-3705464.gif"/></alternatives></inline-formula>, and <inline-formula><tex-math notation="LaTeX">$1.27\times$</tex-math><alternatives><mml:math><mml:mrow><mml:mn>1</mml:mn><mml:mo>.</mml:mo><mml:mn>27</mml:mn><mml:mo>×</mml:mo></mml:mrow></mml:math><inline-graphic xlink:href="xia-ieq5-3705464.gif"/></alternatives></inline-formula> on average.

Ziqi Wang, Xiaoyu Xia, Ibrahim Khalil et al. · 0 citations
#edge computing Sep 2026

D3NN: Adaptive Partitioning and Cross-Tier Resource Orchestration for Cloud–Edge Collaborative Inference

Deep Neural Networks (DNNs) have become foundational to intelligent systems, yet deploying them efficiently under strict latency, resource, and privacy constraints remains challenging. While cloud-only inference suffers from transmission latency and privacy risks, and edge-only execution is limited by hardware capacity, cloud–edge collaborative inference offers a practical middle ground by combining the cloud’s compute strength with the edge’s proximity to data sources for low-latency, scalable, and privacy-aware inference. However, realizing this potential requires adaptive DNN partitioning that responds to dynamic workloads and network conditions, as well as fine-grained cross-tier resource orchestration to avoid bottlenecks and ensure system stability. To this end, we propose DDPG-DRPA-driven Deep Neural Network(D3NN), a novel and efficient framework for partitioned DNN deployment across cloud and edge resources. We formulate the pipeline partitioning of DNNs as a Markov Decision Process (MDP). A value function is trained using the Deep Deterministic Policy Gradient (DDPG) algorithm, and a Dynamic Resource Partitioning Agent (DRPA) allocates suitable cloud or edge resources to each DNN layer according to specific task types. As a result, D3NN adapts dynamically to both environmental conditions and task requirements. Under maximum task arrival rate scenarios, our approach reduces inference latency by 13.7% compared to pure cloud-based inference and by 33.5% compared to pure edge-based inference, demonstrating its practical effectiveness in resource-constrained cloud–edge systems.

Yong Zhao, Zhenjia Mo, Qiang He et al. · 0 citations
#edge computing Sep 2026

Toward 6G Edge Intelligence: Lightweight LLMs for Intent-Driven Network Automation

Future 6G networks are envisaged to tightly integrate communication, sensing, and computing, demanding real-time, intent-driven intelligence at the edge. While large language models (LLMs) excel in intent recognition and semantic reasoning, their application to real-time network lifecycle management at the edge is limited by heterogeneous application intents (APPIs), dynamic network conditions, and severe resource constraints. This paper proposes a novel lightweight LLM architecture, KGLlama-KD, that synergizes knowledge graphs (KGs) with knowledge distillation (KD) to enable intent-driven networking and enhance 6G edge intelligence. Specifically, a KG is constructed to formally describe the relationships among application scenarios, functional primitives, performance requirements within APPIs, and the correspondences between APPIs and network service requests (NSRs), thereby producing a structured intent training dataset. Building upon the Llama 3 foundation model, a two-phase optimization framework is designed to support lightweight edge deployment while preserving translation fidelity. The LLM is first fine-tuned with KG guidance and compressed via KD in the cloud, and then deployed on resource-constrained edge nodes to perform real-time, accurate, and efficient APPIs interpretation. Experiments validate that KGLlama-KD achieves 95% accuracy for APPI understanding, surpassing DeepSeek and Qwen by an average of 8%. The distilled model reduces inference latency by 60% compared to full-scale LLMs, fulfilling the sub-100 ms requirement for 6G latency-sensitive services.

Bing Wu, Sai Zou, Minghui Liwang et al. · 4 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.