Graph Convolutional Multi-Agent Reinforcement Learning for Cooperative UAV Swarm Navigation
Abstract
Cooperative navigation of unmanned aerial vehicle (UAV) swarms in obstacle-rich environments requires simultaneous goal reaching, collision avoidance, and decentralized coordination. This paper presents a graph convolutional network-based multi-agent reinforcement learning (GCN-MARL) framework that explicitly represents the time-varying communication topology of the UAV swarm as a dynamic attributed graph. Neighboring UAV states are aggregated through graph convolution to incorporate local inter-agent interactions into cooperative decision-making. The policy is trained using Multi-Agent Proximal Policy Optimization (MAPPO) under a centralized training with decentralized execution (CTDE) paradigm, together with a composite reward that accounts for goal-directed motion, safety, and communication connectivity. In a five-UAV simulation environment, the proposed GCN-MARL method achieves a 92.0% mission success rate and an 11.0% collision rate, compared with 73.0% and 28.0%, respectively, for an independent PPO baseline. The proposed method also reduces the average completion steps and path length. These results demonstrate the effectiveness of incorporating dynamic topology-aware neighborhood information into cooperative UAV navigation and provide a graph-based framework for decentralized coordination in autonomous multi-UAV systems.