2026· IEEE Transactions on Cognitive Communications and Networking· Vol 12, pp. 9917-9934· 0 citations· 36 references
Computer Science
Abstract
Multi-uncrewed aerial vehicle (UAV) cooperative mobile edge computing (MEC) systems present significant challenges owing to task causal dependencies, dynamic channel variations, and multi-dimensional resource coupling. In this study, a multi-UAV cooperative MEC system with task causality constraints and time-varying wireless channels is considered, and the joint optimization of task offloading, task migration, dynamic UAV clustering, and continuous UAV trajectory planning is investigated. The objective is to minimize the long-term weighted sum of system latency and energy consumption while ensuring task queue stability. Lyapunov optimization is first introduced to transform the formulated stochastic mixed-integer problem into a deterministic per-slot optimization. Afterward, a dual-timescale graph-enhanced multi-agent proximal policy optimization (DT-HGMAPPO) framework is proposed to coordinate long-timescale UAV clustering and trajectory planning with short-timescale task offloading and migration. Specifically, this framework decouples the problem by utilizing dual-layer weighted hypergraph matching (DL-WHM) for joint clustering and association, a dynamic priority scoring (DPS) mechanism for intra-cluster load balancing, and a graph-enhanced MAPPO algorithm for trajectory optimization. Simulation results reveal that the proposed DT-HGMAPPO algorithm outperforms conventional multi-agent deep reinforcement learning baselines in terms of both convergence speed and policy stability. It achieves a final reward that is at least 18.0% greater than that of other multi-agent algorithms. Moreover, the proposed framework reduces total system cost by 33.3% compared with MASAC and 18.9% compared with MATD3, thereby achieving a superior delay-energy trade-off while ensuring queue stability.
A Lyapunov-based joint optimization framework for UAV-enabled MEC systems achieves a balanced tradeoff between delay, energy consumption, and UAV flight activity, supporting energy-efficient and delay-aware UAV-MEC operation.
Lei Li, Xue Gao, Quansheng Guan· Electronics· 0 citations
This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.
Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al.· Journal of King Saud Univers...· 0 citations
A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner and outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.
Tiankui Zhang, Wenlong Xu, Tianyi Shi et al.· IEEE Internet of Things Jour...· 0 citations
The Sequentially Extended Consensus-Based Bundle Algorithm (SECBBA), a deadlock-free distributed scheduling framework, is proposed, extending through the integration of a deadlock detection and resolution mechanism based on directed graph Depth-First Search, thereby guaranteeing conflict-free task allocation.
A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.
M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al.· 0 citations
Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.
Yeguang Qin, Jie Tang, Fengxiao Tang et al.· IEEE Transactions on Communi...· 0 citations