Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 56 references
TL;DR
This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.
Abstract
Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) is a key enabler for meeting the stringent low-latency and energy-efficiency requirements of emerging low-altitude economy applications. However, achieving these objectives remains challenging due to dynamic environments, limited communication and computation resources, and the heterogeneity of network entities. This paper investigates the long-term joint optimization framework that minimizes system-wide latency and energy consumption simultaneously by coordinating UAV association, subchannel selection, uplink/downlink power allocation, and computational resource distribution. This sequential decision-making process is formulated into a partially observable Markov decision process (POMDP) to account for localized observations and dynamic channel states. To solve it, we propose a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices (UDs) and UAVs act as heterogeneous agents. This architecture utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers. Numerical results demonstrate that the proposed scheme effectively navigates the high-dimensional action space and achieves superior convergence and cost reduction compared to benchmarks, including PPO, independent PPO (iPPO), Q-learning multi-agent extension (QMIX), value decomposition networks (VDN), independent deep Q-network (iDQN), and genetic algorithm (GA).
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.
Yongbin Wang, Peng Lin, Yan Liu et al.· IEEE Wireless Communications...· 0 citations
Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.
Meng Yi, V. Lee, Miao Du et al.· IEEE Transactions on Cogniti...· 0 citations
A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner and outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.
Tiankui Zhang, Wenlong Xu, Tianyi Shi et al.· IEEE Internet of Things Jour...· 0 citations
In satellite-UAV assisted communication networks, jointly optimizing the UAV’s trajectory and the multi-agent scheduling decisions to minimize the age of information (AoI) is a notoriously challenging problem. The complexity is compounded by the fundamental heterogeneity between the satellite and UAV agents, including their disparate action spaces, partial observations, and differing energy-consumption and communication-cost penalties. To address this, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose a novel heterogeneous multi-agent compound-action proximal policy optimization (HMACPPO) algorithm. HMACPPO leverages a centralized training with decentralized execution (CTDE) framework, using role-specific decentralized actors together with agent-specific centralized critics conditioned on the global state. Specifically, the UAV employs a compound PPO (CPPO) actor for its hybrid action space, while the satellite uses a PPO actor for discrete scheduling. Extensive simulations show that HMACPPO outperforms the compared baselines, and that the resulting coordinated policy effectively manages the trade-off between AoI, UAV energy consumption, and operational cost.
Weijie Zhou, Mengjie Yi, Yan Zhang et al.· IEEE Transactions on Cogniti...· 0 citations
Driven by the vision of a thriving low-altitude economy and aiming to provide on-demand services for diverse entities, this paper investigates an integrated sensing and communication (ISAC)-enabled low-altitude wireless network (LAWN). Benefiting from flexible mobility and cost-effective cooperative deployment, multiple ISAC-enabled uncrewed aerial vehicles (UAVs) are emerging as an ISAC paradigm for on-demand deployment in LAWN. However, due to the complex inter-UAV interference and resource coupling in LAWN, it is difficult to properly coordinate different constrained resources, including spatial deployment, energy, and wireless channels, to simultaneously meet the sensing and communication requirements. To address these challenges, this paper formulates a sensing–communication optimization (SCO) problem in LAWN by jointly optimizing subcarrier allocation, transmit power allocation, and three-dimensional (3D) UAV deployments to maximize network utility while satisfying quality of service (QoS) requirements for multiple users and target sensing mutual information (MI) requirements. To enable efficient solutions, we propose a hierarchical optimization approach that vertically decouples the SCO problem into two subproblems: a top level employing a Gibbs Sampling–based multi-UAV 3D deployment algorithm for efficient exploration and deployment optimization, and a bottom level performing resource allocation via a dual-based joint power and subcarrier allocation algorithm. Simulation results demonstrate that the proposed approach achieves a favorable trade-off between communication and sensing and significantly enhances the overall performance and adaptability of the LAWN.
Cheng Ma, Zewei Jing, Qinghai Yang et al.· IEEE Transactions on Wireles...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.