Skip to content
Conference

Joint Trajectory and Scheduling Optimization for UAV-Assisted 6G Networks: A Deep Reinforcement Learning Approach with Throughput–AoI Trade-off

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 118-121 · 0 citations · 14 references

Abstract

Unmanned aerial vehicle (UAV) communications are a promising enabler for 6G networks, offering flexible deployment and strong line-of-sight channel conditions. Effective UAV operation requires jointly optimizing trajectory and user scheduling to balance throughput and information freshness. This paper proposes a proximal policy optimization (PPO)-based deep reinforcement learning (DRL) framework that controls UAV movement and user scheduling together via a joint MultiDiscrete action space. We formulate a Markov decision process for a 8-user, $1000 \times 1000 \mathrm{~m}^{2}$ service area with a 3GPP TR 36.777-compliant channel model, where the agent selects both its next position and which user to serve at each time slot. The proposed PPO policy achieves 85.75 Mbps mean throughput, a 24.4% improvement over the AoI-greedy baseline, while reducing mean AoI by 87.7% compared to the throughput-greedy baseline, reaching a Pareto-optimal trade-off between the two competing objectives. An ablation study over the AoI penalty weight confirms a clear throughput-AoI trade-off, validating the joint design.

View source

Similar papers

Conference Jul 2026

Reinforcement Learning-Based Decode-and-Forward UAV Relay Trajectory Optimization

Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.

Aniket Subbanwar, Ojas Joshi, Amit Agarwal · 0 citations
Open access 2026

LLM-Guided Multi-Agent Joint Velocity and Spectrum Optimization in Advanced Air Mobility

In Advanced Air Mobility (AAM) applications, jointly optimizing multi-agent motion control along predefined flight routes and spectrum access is highly challenging due to the tight coupling among mobility, interference, and safety constraints under limited spectrum resources. This paper proposes a Large Language Model (LLM)-guided cooperative decision-making framework for joint velocity control and bidirectional channel selection in an AAM system with Aerial Vehicles (AVs) communicating with ground Base Stations (BSs) while following predefined linear routes. We formulate the problem as a cooperative Markov game with a discrete action space that includes both velocity selection and uplink and downlink channel access, while satisfying Signal-to-Interference-plus-Noise Ratio (SINR) quality requirements and collision avoidance constraints. To obtain reliable expert behavior, we first learn a near-optimal policy using Multi-Agent Reinforcement Learning (MARL) with Value Decomposition Dueling Double Deep Q-Networks (VD3QN). We then treat joint decision-making as a sequence generation task and employ Large Language Models (LLMs) to generate complete joint action sequences from structured environment descriptions, under both Prompt Engineering (PE) and Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) on expert demonstrations. Extensive simulations show that structured prompting improves decision quality, while LoRA fine-tuning further increases reward, reduces variance, and yields decisions that closely match the expert policy. Beyond this imitation role, the LLM layer turns 6-AV expert demonstrations into a sequence-level decision generator. In an unseen 10-AV scenario, this generator achieves stronger zero-shot generalization than the VD3QN policy transferred from the 6-AV environment.

Qingyang Li, Adnan Quadri, Hongxiang Li et al. · 0 citations
Open access Jul 2026

Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture

A multi-agent deep reinforcement learning (MADRL) methodology based on the Multi-Agent Proximal Policy Optimization (MAPPO) approach, which simultaneously achieves high field coverage completeness, robust communication energy efficiency, and a high depot-return rate under hard battery constraints without any inter-UAV communication overhead at execution time.

Ayman Massaoudi, Walid Aydi · 0 citations
2026

Optimization of 3-D Trajectory and Resource Allocation in Multi-UAV Communications Under a Probabilistic Channel Model

This study investigates the optimization of three-dimensional (3D) trajectory planning and resource allocation in unmanned aerial vehicle (UAV)-enabled wireless networks with no-fly zones (NFZs) using a deep learning framework. The objective is to maximize the minimum average spectral efficiency (SE) among mobile users served by multiple UAVs while addressing key challenges, including interference from concurrent UAV transmissions, collision avoidance, and NFZ constraints. A realistic probabilistic channel model is considered, where the likelihood of a line-of-sight (LoS) condition is modeled as a function of the elevation angle in the air-to-ground (A2G) link. To solve the formulated optimization problem, a novel deep learning framework with specialized deep neural network (DNN) structures is developed. This framework jointly optimizes 3D UAV trajectory planning and resource allocation, employing an unsupervised learning-based training approach that eliminates the need for labeled data. Performance evaluations demonstrate that the proposed scheme effectively accounts for the probabilistic channel model and co-channel interference while accounting for collision avoidance and NFZ-related constraints. Moreover, it outperforms baseline methods by achieving a higher minimum average SE with real-time computational efficiency, making it practical for UAV-assisted wireless networks.

Woongsup Lee, Howon Lee, Kisong Lee · 0 citations
2026

Collaborative Trajectory and Resource Optimization in Multi-UAV MEC Under Jamming: An LLM-Guided MARL Framework

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.

Yeguang Qin, Jie Tang, Fengxiao Tang et al. · 0 citations
Conference Jul 2026

Double Deep Reinforcement Learning–Based UAV Positioning for Throughput Optimization in Wireless Networks

This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.

Berke Kilinç, M. Ö. Efe · 0 citations