Skip to content

Multi-UAV-Aided Data Collection in Complex 3-D Urban Environments: A MADRL Approach Enhanced With Pheromone-Reward Shaping

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 9883-9901 · 0 citations · 51 references
Computer Science

Abstract

This paper investigates the problem of cooperative multiple unmanned aerial vehicles (UAVs) data collection for Internet of Things (IoT) networks in dense urban environments. Unlike existing studies that predominantly rely on idealized spatial models and average-based probabilistic channel models, this work explicitly accounts for realistic 3-D building distributions and deterministically models ground-to-air (G2A) channel blockages. We formulate a joint optimization problem to minimize the total task completion time, subject to stringent system throughput, flight dynamics, and energy constraints. To tackle the highly coupled challenges of node scheduling and trajectory planning, we propose a lightweight two-stage heuristic strategy for dynamic access control, along with a multi-agent reinforcement learning for trajectory planning. Crucially, to overcome the severe sparse-reward bottleneck inherent in complex 3-D obstacle avoidance, we introduce a Pheromone-based Reward Shaping (PRS) mechanism. By mathematically integrating the UAV’s kinematic state with deterministic environmental feedback, PRS effectively transforms the sparse-reward navigation challenge into a dense and smooth gradient, thereby profoundly accelerating policy convergence. Extensive simulations demonstrate that the proposed MATD3-PRS framework significantly outperforms representative baselines, achieving superior performance in task completion time, flight trajectory efficiency, and overall energy saving.

View source

Similar papers

2026

Optimization of 3-D Trajectory and Resource Allocation in Multi-UAV Communications Under a Probabilistic Channel Model

This study investigates the optimization of three-dimensional (3D) trajectory planning and resource allocation in unmanned aerial vehicle (UAV)-enabled wireless networks with no-fly zones (NFZs) using a deep learning framework. The objective is to maximize the minimum average spectral efficiency (SE) among mobile users served by multiple UAVs while addressing key challenges, including interference from concurrent UAV transmissions, collision avoidance, and NFZ constraints. A realistic probabilistic channel model is considered, where the likelihood of a line-of-sight (LoS) condition is modeled as a function of the elevation angle in the air-to-ground (A2G) link. To solve the formulated optimization problem, a novel deep learning framework with specialized deep neural network (DNN) structures is developed. This framework jointly optimizes 3D UAV trajectory planning and resource allocation, employing an unsupervised learning-based training approach that eliminates the need for labeled data. Performance evaluations demonstrate that the proposed scheme effectively accounts for the probabilistic channel model and co-channel interference while accounting for collision avoidance and NFZ-related constraints. Moreover, it outperforms baseline methods by achieving a higher minimum average SE with real-time computational efficiency, making it practical for UAV-assisted wireless networks.

Woongsup Lee, Howon Lee, Kisong Lee · 0 citations
Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
2026

Energy-Aware Multi-UAV Collaboration for Data Collection and Trajectory Planning With MADDPG

Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.

Jing Mei, Jinglei Xu, Zhao Tong et al. · 0 citations
Open access Jul 2026

Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture

A multi-agent deep reinforcement learning (MADRL) methodology based on the Multi-Agent Proximal Policy Optimization (MAPPO) approach, which simultaneously achieves high field coverage completeness, robust communication energy efficiency, and a high depot-return rate under hard battery constraints without any inter-UAV communication overhead at execution time.

Ayman Massaoudi, Walid Aydi · 0 citations
2026

Collaborative Trajectory and Resource Optimization in Multi-UAV MEC Under Jamming: An LLM-Guided MARL Framework

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.

Yeguang Qin, Jie Tang, Fengxiao Tang et al. · 0 citations
#edge computing Open access Aug 2026

Collaborative resource allocation in UAV-assisted MEC networks: A heterogeneous MAPPO scheme

This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.

Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al. · 0 citations