2026· IEEE Transactions on Cognitive Communications and Networking· Vol 12, pp. 9992-10005· 0 citations· 33 references
Abstract
The evolution of uncrewed aerial vehicles (UAVs) into embodied intelligent agents in the low-altitude economy is reshaping edge computing networks. However, the high mobility of UAVs induces severe topology dynamics, limiting the efficacy of traditional fully connected multiagent reinforcement learning because of dimensionality and credit assignment challenges. Furthermore, existing graph attention approaches neglect explicit communication boundary constraints, leading to mismatches between value evaluation and physical topology. To address these challenges, this paper proposes the spatial-aware graph attention multiagent twin delayed deep deterministic policy gradient (SAGA-MATD3) algorithm. By embedding a dynamic spatial masking mechanism based on the communication radius into the critic network, the proposed method enforces physical reachability constraints and attenuates extraneous noise. Simulation results demonstrate that SAGA-MATD3 significantly reduces service latency and improves fairness under an acceptable energy-consumption tradeoff, achieving a 37.1% improvement in convergence reward and enabling the self-organization of robust load-balanced mesh topologies.
A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.
M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al.· 0 citations
This paper proposes the Locally-decoupled and Embedding-enhanced Multi-Agent Deep Deterministic Policy Gradient (LDE-MADDPG) algorithm to address poor scalability and delayed response in drone swarm dynamic obstacle avoidance under complex cooperative environments. Such autonomous coordination capabilities are also important for distributed sensing, wireless networking, and electromagnetic information exchange in future intelligent aerial systems. The algorithm introduces three key innovations beyond standard MADDPG: a Graph Attention Network module that encodes variable-length observations into fixed-dimensional embeddings for swarm-size generalization; a dual-path critic with a global branch guiding policy updates and a local branch specializing in obstacle avoidance evaluation; and a hierarchical reward integrating multi-objective signals. Evaluated across eight static and dynamic obstacle scenarios, LDE-MADDPG achieves significantly lower collision rates (2.1%–4.2% in static scenarios and 3.8%–7.2% in dynamic scenarios) than state-of-the-art baselines and reaches a 97.5% mission completion rate in 100 random scenarios. The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
X. Fang, K. Chen, C. Ren et al.· Advanced Electromagnetics· 0 citations
While deploying hierarchical vision models to process mission-critical tasks, UAV edge systems must adaptively update the models to sustain inference reliability under low-level environmental corruption. However, existing work has overlooked the optimal timing for model updates, the impracticality of relying on real-time expert labels, and the significant bandwidth and energy constraints of UAVs. This paper proposes a joint model update scheduling and resource allocation framework, aiming to maximize long-term semantic fidelity and resource efficiency of UAV edge intelligence systems. To address the challenge of label-free semantic evaluation, we formulate the Online Semantic Disagreement Rate (OSDR) as a proxy for timely update triggering, thereby enabling fine-grained Sensitivity-Aware Structural Synchronization (SASS). Furthermore, to overcome the curse of dimensionality in hybrid action spaces and effectively bound long-term energy budgets, we propose a Lyapunov-guided discrete reinforcement learning algorithm that performs action space dimensionality reduction and transforms constraints into virtual queue stability problems. The reported experimental results, based on real traffic traces, demonstrate that the proposed framework consistently outperforms representative baselines in semantic recovery efficiency and update triggering precision, by satisfying long-term energy budget and by reducing average risk backlog by up to 33.3\% in the dynamic environmental corruption scenario.
Feng He, Alireza Furutanpey, Paolo Bellavista et al.· 0 citations
Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.
Meng Yi, V. Lee, Miao Du et al.· IEEE Transactions on Cogniti...· 0 citations
Safe-separation-and-collision-avoidance unmanned aerial vehicle (UAV) swarms are increasingly used for inspection, emergency response, environmental monitoring, and search-and-rescue support in cluttered airspace where communication links may be delayed, degraded, or intermittently unavailable. These applications require heterogeneous vehicles to maintain situational awareness, allocate tasks, and avoid hazards under partial observability and changing team topology. To address these challenges, this paper proposes a Hierarchical Graph-Attention Multi-Agent Reinforcement Learning architecture (HG-MARL) for safe-separation-and-collision-avoidance heterogeneous UAV swarm coordination. The proposed framework decomposes the task into high-level resource allocation and low-level local-control execution, uses graph attention for changing swarm topology, and applies Transformer memory, action masking, potential-field reward shaping, and domain-randomized simulation training. In the multi-scenario simulation summaries, HG-MARL achieves 92.9%, 89.8%, and 82.6% task success in Scenarios A–C, respectively, improving upon MAPPO by 15.1, 21.4, and 20.1 percentage points. Summary-statistic Welch tests show that all six HG-MARL comparisons against MAPPO and QMIX yield p<0.01 with large effect sizes. Fair-control, reward-sensitivity, communication-degradation, safety-ablation, training-stability, latency, and transfer-oriented stress tests further support the contributions of the integrated architecture. The validation scope is simulator-based, with platform-level flight/HIL evaluation discussed as future work. These results suggest that HG-MARL is a promising simulation-validated framework for civilian UAV swarm coordination in collision-and-separation-critical and communication-degraded environments.
Xudong Zhang, Junqiang Bai, K. Chen et al.· Drones· 0 citations