Skip to content
#edge computing Open access

Collaborative resource allocation in UAV-assisted MEC networks: A heterogeneous MAPPO scheme

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 56 references

TL;DR

This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.

Abstract

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) is a key enabler for meeting the stringent low-latency and energy-efficiency requirements of emerging low-altitude economy applications. However, achieving these objectives remains challenging due to dynamic environments, limited communication and computation resources, and the heterogeneity of network entities. This paper investigates the long-term joint optimization framework that minimizes system-wide latency and energy consumption simultaneously by coordinating UAV association, subchannel selection, uplink/downlink power allocation, and computational resource distribution. This sequential decision-making process is formulated into a partially observable Markov decision process (POMDP) to account for localized observations and dynamic channel states. To solve it, we propose a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices (UDs) and UAVs act as heterogeneous agents. This architecture utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers. Numerical results demonstrate that the proposed scheme effectively navigates the high-dimensional action space and achieves superior convergence and cost reduction compared to benchmarks, including PPO, independent PPO (iPPO), Q-learning multi-agent extension (QMIX), value decomposition networks (VDN), independent deep Q-network (iDQN), and genetic algorithm (GA).

Read PDF

Similar papers

Open access Jul 2026

Toward Low-Delay and Energy-Efficient UAV-Assisted MEC Systems Through Intelligent Resource Allocation

A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.

W. Saber, Hanan Algamil, Fifi Farouk et al. · 0 citations
2026

Energy-Efficient Task Offloading and Load Balancing for Multi-UAV-Assisted Vehicular Networks

The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.

Yongbin Wang, Peng Lin, Yan Liu et al. · 0 citations
2026

SkySched: A Hierarchical and Scalable Reinforcement Learning Framework for Multi-UAV Vehicular Edge Computing Network

Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.

Meng Yi, V. Lee, Miao Du et al. · 0 citations
#edge computing Open access Aug 2026

Distributed Trajectory Planning and Resource Allocation for Dynamic Multi-UAV Collaborative Computing

A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner and outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.

Tiankui Zhang, Wenlong Xu, Tianyi Shi et al. · 0 citations
2026

Optimizing Information Freshness in Satellite-UAV IoRT Networks: A Heterogeneous Multi-Agent Approach

In satellite-UAV assisted communication networks, jointly optimizing the UAV’s trajectory and the multi-agent scheduling decisions to minimize the age of information (AoI) is a notoriously challenging problem. The complexity is compounded by the fundamental heterogeneity between the satellite and UAV agents, including their disparate action spaces, partial observations, and differing energy-consumption and communication-cost penalties. To address this, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose a novel heterogeneous multi-agent compound-action proximal policy optimization (HMACPPO) algorithm. HMACPPO leverages a centralized training with decentralized execution (CTDE) framework, using role-specific decentralized actors together with agent-specific centralized critics conditioned on the global state. Specifically, the UAV employs a compound PPO (CPPO) actor for its hybrid action space, while the satellite uses a PPO actor for discrete scheduling. Extensive simulations show that HMACPPO outperforms the compared baselines, and that the resulting coordinated policy effectively manages the trade-off between AoI, UAV energy consumption, and operational cost.

Weijie Zhou, Mengjie Yi, Yan Zhang et al. · 0 citations
2026

Hierarchical Optimization of UAV Deployment and Resource Allocation for ISAC-Enabled Low-Altitude Wireless Networks

Driven by the vision of a thriving low-altitude economy and aiming to provide on-demand services for diverse entities, this paper investigates an integrated sensing and communication (ISAC)-enabled low-altitude wireless network (LAWN). Benefiting from flexible mobility and cost-effective cooperative deployment, multiple ISAC-enabled uncrewed aerial vehicles (UAVs) are emerging as an ISAC paradigm for on-demand deployment in LAWN. However, due to the complex inter-UAV interference and resource coupling in LAWN, it is difficult to properly coordinate different constrained resources, including spatial deployment, energy, and wireless channels, to simultaneously meet the sensing and communication requirements. To address these challenges, this paper formulates a sensing–communication optimization (SCO) problem in LAWN by jointly optimizing subcarrier allocation, transmit power allocation, and three-dimensional (3D) UAV deployments to maximize network utility while satisfying quality of service (QoS) requirements for multiple users and target sensing mutual information (MI) requirements. To enable efficient solutions, we propose a hierarchical optimization approach that vertically decouples the SCO problem into two subproblems: a top level employing a Gibbs Sampling–based multi-UAV 3D deployment algorithm for efficient exploration and deployment optimization, and a bottom level performing resource allocation via a dual-based joint power and subcarrier allocation algorithm. Simulation results demonstrate that the proposed approach achieves a favorable trade-off between communication and sensing and significantly enhances the overall performance and adaptability of the LAWN.

Cheng Ma, Zewei Jing, Qinghai Yang et al. · 0 citations

Related blog posts

Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.