Skip to content

Optimization of 3-D Trajectory and Resource Allocation in Multi-UAV Communications Under a Probabilistic Channel Model

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 10794-10809 · 0 citations · 51 references

Abstract

This study investigates the optimization of three-dimensional (3D) trajectory planning and resource allocation in unmanned aerial vehicle (UAV)-enabled wireless networks with no-fly zones (NFZs) using a deep learning framework. The objective is to maximize the minimum average spectral efficiency (SE) among mobile users served by multiple UAVs while addressing key challenges, including interference from concurrent UAV transmissions, collision avoidance, and NFZ constraints. A realistic probabilistic channel model is considered, where the likelihood of a line-of-sight (LoS) condition is modeled as a function of the elevation angle in the air-to-ground (A2G) link. To solve the formulated optimization problem, a novel deep learning framework with specialized deep neural network (DNN) structures is developed. This framework jointly optimizes 3D UAV trajectory planning and resource allocation, employing an unsupervised learning-based training approach that eliminates the need for labeled data. Performance evaluations demonstrate that the proposed scheme effectively accounts for the probabilistic channel model and co-channel interference while accounting for collision avoidance and NFZ-related constraints. Moreover, it outperforms baseline methods by achieving a higher minimum average SE with real-time computational efficiency, making it practical for UAV-assisted wireless networks.

View source

Similar papers

2026

Multi-UAV-Aided Data Collection in Complex 3-D Urban Environments: A MADRL Approach Enhanced With Pheromone-Reward Shaping

This paper investigates the problem of cooperative multiple unmanned aerial vehicles (UAVs) data collection for Internet of Things (IoT) networks in dense urban environments. Unlike existing studies that predominantly rely on idealized spatial models and average-based probabilistic channel models, this work explicitly accounts for realistic 3-D building distributions and deterministically models ground-to-air (G2A) channel blockages. We formulate a joint optimization problem to minimize the total task completion time, subject to stringent system throughput, flight dynamics, and energy constraints. To tackle the highly coupled challenges of node scheduling and trajectory planning, we propose a lightweight two-stage heuristic strategy for dynamic access control, along with a multi-agent reinforcement learning for trajectory planning. Crucially, to overcome the severe sparse-reward bottleneck inherent in complex 3-D obstacle avoidance, we introduce a Pheromone-based Reward Shaping (PRS) mechanism. By mathematically integrating the UAV’s kinematic state with deterministic environmental feedback, PRS effectively transforms the sparse-reward navigation challenge into a dense and smooth gradient, thereby profoundly accelerating policy convergence. Extensive simulations demonstrate that the proposed MATD3-PRS framework significantly outperforms representative baselines, achieving superior performance in task completion time, flight trajectory efficiency, and overall energy saving.

Haitao Chen, Xinfeng Deng, Zhe Wang et al. · 0 citations
Jul 2026

Joint optimization of 3D deployment and power allocation for multi-UAV base stations

In temporary emergency communication coverage scenarios where terrestrial communication infrastructure is damaged or lacks sufficient capacity, UAVs equipped with base stations have emerged as an effective solution due to their flexible deployment and rapid response capability. However, in multi-UAV networks, the three-dimensional deployment of UAVs significantly affects air-to-ground link quality, while power allocation further determines the level of system interference and throughput performance. To address this issue, this paper considers a multi-UAV communication system and jointly takes into account user link reliability and service requirement satisfaction, thereby establishing a joint optimization model for QoS-constrained coverage and network throughput. To address the non-convex joint optimization problem, a problem-tailored dual-population cooperative NSGA-II framework, termed IDPC-NSGA-II, is developed. By coupling dual-population evolution, adaptive mutation, uncovered-user-guided local search, and interference-aware repair with the characteristics of multi-UAV emergency communications, the proposed method improves the trade-off between QoS-constrained coverage and network throughput. Simulation results in a representative emergency communication scenario show that the proposed method achieves a favorable trade-off between QoS-constrained coverage and throughput, and outperforms the compared algorithms under the considered network setting.

Guifen Chen, Ruiyang Liu · 0 citations
2026

3-D Trajectory Design Based on Deep Reinforcement Learning for UAV-Assisted Communication Networks

Most of the existing UAV-assisted communication networks provide service only for static users or deterministically moving ones. In fact, for some complex and dynamically changing scenarios, the users communicating to the UAV may move randomly, with unpredictable mobility. The uncertainty of users’ movements poses a challenge to the guarantee of stable network performance. To tackle this, the paper investigates a UAV-assisted communication network, where a UAV provides communication service for ground users which are moving randomly. We collectively factor in ground user mobility, task duration, and UAV flight restrictions to design precise 3D trajectory for UAV, and formulate them into an optimization problem, aiming to maximize the network throughput while minimizing UAV energy consumption. Considering the dynamics caused by users’ uncertain movement, we transform the optimization problem into a Markov decision process (MDP), then improve the twin-delayed deep deterministic policy gradient (TD3) to design UAV’s 3D trajectory. By utilizing the prior knowledge to accelerate the exploration efficiency, we propose a trajectory design algorithm based on prior knowledge-TD3 (PKTD3-TD), enabling UAV to autonomously adjust flight parameters by leveraging environmental observations under dynamic conditions for enhancing flexibility and intelligence. Simulation results show that our proposed scheme outperforms the compared ones in terms of communication link quality, network throughput and UAV’s energy consumption.

Min Li, M. Dong, Hong Wang et al. · 0 citations
#edge computing Open access Aug 2026

Distributed Trajectory Planning and Resource Allocation for Dynamic Multi-UAV Collaborative Computing

A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner and outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.

Tiankui Zhang, Wenlong Xu, Tianyi Shi et al. · 0 citations
Conference Jul 2026

Joint Trajectory and Scheduling Optimization for UAV-Assisted 6G Networks: A Deep Reinforcement Learning Approach with Throughput–AoI Trade-off

Unmanned aerial vehicle (UAV) communications are a promising enabler for 6G networks, offering flexible deployment and strong line-of-sight channel conditions. Effective UAV operation requires jointly optimizing trajectory and user scheduling to balance throughput and information freshness. This paper proposes a proximal policy optimization (PPO)-based deep reinforcement learning (DRL) framework that controls UAV movement and user scheduling together via a joint MultiDiscrete action space. We formulate a Markov decision process for a 8-user, $1000 \times 1000 \mathrm{~m}^{2}$ service area with a 3GPP TR 36.777-compliant channel model, where the agent selects both its next position and which user to serve at each time slot. The proposed PPO policy achieves 85.75 Mbps mean throughput, a 24.4% improvement over the AoI-greedy baseline, while reducing mean AoI by 87.7% compared to the throughput-greedy baseline, reaching a Pareto-optimal trade-off between the two competing objectives. An ablation study over the AoI penalty weight confirms a clear throughput-AoI trade-off, validating the joint design.

Quang Tuan Do, Tung Son Do, Thanh Phung Truong et al. · 0 citations