Deep Reinforcement Learning with Lagrangian-Guided Reward Design for Joint Clustering and Spectrum Allocation in Dynamic UAV Swarms
Abstract
Joint clustering and spectrum allocation in dynamic unmanned aerial vehicle (UAV) swarms is a coupled integer programming problem whose complexity increases rapidly with the number of UAVs. In dynamic communication scenarios with time-varying topology, channel conditions, and traffic demands, conventional optimization methods are difficult to apply in real time. Although deep reinforcement learning (DRL) is promising for online decision-making, its performance is often limited by the inconsistency between the original constrained optimization objective and the reward used in training. To address this issue, this paper proposes a DRL method with Lagrangian-guided reward design for joint clustering and spectrum allocation in dynamic UAV swarms. Lagrangian relaxation is incorporated into the reward formulation to better align policy learning with the original constrained optimization problem. In addition, an action mapping mechanism is introduced to generate feasible clustering and spectrum allocation decisions under cluster-size constraints. Simulation results show that the proposed method improves system capacity, task completion performance, and delay performance, while achieving faster and more stable convergence than baseline reinforcement learning methods.