A graph-enhanced centralized critic is proposed that injects topology-aware relational embeddings into value estimation while leaving decentralized actors unchanged, with centralized-training overhead and dense 16-secondary-user settings as practical limitations.
Abstract
Cooperative spectrum access in multi-UAV cognitive radio networks requires decentralized control under dynamic primary-user activity, interference coupling, and partial observability. Standard multi-agent deep deterministic policy gradient (MADDPG) follows centralized training with decentralized execution, but its centralized critic usually processes the joint state-action context as a flat vector and therefore does not explicitly encode inter-UAV topology. This paper proposes a graph-enhanced centralized critic that injects topology-aware relational embeddings into value estimation while leaving decentralized actors unchanged. Under a unified eight-seed dynamic-spectrum protocol, the weighted radial basis function (weighted-RBF) graph critic improves mean reward from 4238.63 to 4346.55 and reduces collision rate from 0.1004 to 0.0383 relative to vanilla MADDPG. A stronger TD3-style multi-agent baseline (MATD3) reaches 4348.47 reward and 0.0495 collision rate, while MATD3 with the weighted-RBF graph critic further improves reward to 4379.29 and lowers collision rate to 0.0346. Additional topology ablation, high-interference, complexity, and scalability analyses show that graph-enhanced critic learning is most useful for collision control and interference-heavy operation, with centralized-training overhead and dense 16-secondary-user settings as practical limitations.
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations
A graph attention network-enhanced multi-agent proximal policy optimization (GAT-MAPPO) framework is proposed for cooperative guidance in adversarial engagement scenarios. A dynamic heterogeneous interaction graph is formulated over interceptors and targets at every decision epoch. Through a multi-head graph attention encoder, relational features capturing both inter-interceptor cooperation and target threat dynamics are adaptively aggregated. These graph-enriched observations are processed by a Centralized-Training, Decentralized-Execution (CTDE) MAPPO architecture, guided by a hierarchical reward function that mandates miss distance minimization, simultaneity of arrival consensus, multi-directional encirclement, and smooth control effort. Furthermore, the integration of a three-stage curriculum learning strategy allows for robust cooperative policy derivation across transitions from rectilinear to highly adaptive evasion patterns, eliminating the need for explicit rule engineering. Extensive Monte Carlo simulations confirm GAT-MAPPO’s superior performance: achieving >95% interception success rate in 4-vs.-4 scenarios and reducing mean simultaneity error by 41.4% compared to the MAPPO baseline. Comprehensive ablation and sensitivity studies validate the critical roles played by graph attention encoding, reward hierarchy design, and progressive curriculum staging.
The evolution of uncrewed aerial vehicles (UAVs) into embodied intelligent agents in the low-altitude economy is reshaping edge computing networks. However, the high mobility of UAVs induces severe topology dynamics, limiting the efficacy of traditional fully connected multiagent reinforcement learning because of dimensionality and credit assignment challenges. Furthermore, existing graph attention approaches neglect explicit communication boundary constraints, leading to mismatches between value evaluation and physical topology. To address these challenges, this paper proposes the spatial-aware graph attention multiagent twin delayed deep deterministic policy gradient (SAGA-MATD3) algorithm. By embedding a dynamic spatial masking mechanism based on the communication radius into the critic network, the proposed method enforces physical reachability constraints and attenuates extraneous noise. Simulation results demonstrate that SAGA-MATD3 significantly reduces service latency and improves fairness under an acceptable energy-consumption tradeoff, achieving a 37.1% improvement in convergence reward and enabling the self-organization of robust load-balanced mesh topologies.
Ye Wang, Jingjing Wang, Jianrui Chen et al.· IEEE Transactions on Cogniti...· 0 citations
Edge-assisted cognitive radio networks require efficient scheduling mechanisms to jointly manage opportunistic spectrum access, task offloading, energy consumption, and latency constraints. Existing multi-agent scheduling approaches often rely on fixed penalty terms or average queue-based constraints, which may not effectively control service-level violations under uncertain spectrum availability and dynamic edge-resource contention. This work proposes SCOPE, a Safe Causal-Graph Primal–Dual multi-agent scheduling framework for energy- and latency-constrained edge-assisted cognitive radio networks. The major strength of SCOPE is its integrated design, where belief-state augmentation improves decision-making under imperfect spectrum sensing, dual-relational causal graph coordination separately models’ interference coupling and computation-resource contention, and CVaR-based primal–dual optimization regulates tail-risk violations of latency and energy constraints. The framework follows a centralized-training and decentralized-execution structure, enabling coordinated learning during training while supporting scalable decentralized scheduling during deployment. Simulation results under dynamic user mobility, stochastic task arrivals, and varying primary-user activity show that SCOPE improves latency, energy efficiency, service-level constraint satisfaction, and throughput compared with existing scheduling methods. Ablation analysis further confirms the individual contribution of belief modeling, graph coordination, and risk-sensitive constraint enforcement.
T. Kannan, M. Lavanya, A. Ponraj et al.· Scientific Reports· 0 citations
6G mobile edge networks are emerging as a key infrastructure for ubiquitous large language model (LLM) inference services. However, conventional edge routing to nearby or well-connected servers falls short for efficient edge LLM inference, as it may miss the user’s KV cache and trigger costly prefill recomputation. To address this challenge, this paper studies an edge inference system assisted by an embodied UAV agent swarm, where UAVs actively sense user mobility and neighboring UAV states to make local decisions on trajectory control, user association, and inference-request routing. The goal is to improve KV-cache reuse while maintaining reliable wireless connectivity, thereby maximizing the system effective token throughput under energy and QoS constraints. We then formulate the joint optimization as a mixed-integer non-linear program and further cast the sequential UAV decision-making process as a decentralized partially observable Markov decision process. To obtain scalable decentralized policies under partial observations, we propose Q-MAA2C, a quantum-enhanced multi-agent advantage actor-critic algorithm for embodied UAV swarm control and inference routing. Q-MAA2C uses quantum actors for local action selection and an entangled split critic for swarm-level value estimation, enabling coordinated policies from partial observations with reduced raw observation exchange. Simulation results indicate that Q-MAA2C yields comparable reinforcement learning rewards to the fully classical baseline while reducing the number of convergence episodes by about 43%. Additionally, the proposed method enhances the system effective token throughput by up to about 134% over other competing methods.
Xiangdong Zheng, Long Luo, Hongfang Yu et al.· IEEE Transactions on Cogniti...· 1 citation
Traffic load balancing is a key near-real-time control function in dense and heterogeneous radio access networks, where uneven and time-varying traffic distributions can severely degrade resource utilisation and quality of service. Within the Open Radio Access Network (O-RAN) architecture, cell individual offset (CIO) control provides an effective mechanism to steer user handovers for load balancing. However, existing rule-based and centralised deep reinforcement learning (DRL) approaches suffer from limited adaptability and poor scalability, due to static assumptions, excessive signalling overhead, and the exponential growth of the joint action space. To address these challenges, this paper proposes a graph-based multi-agent reinforcement learning (MARL) framework with centralised training and distributed execution for CIO-based load balancing in O-RAN. Distributed actors perform decentralised near-real-time CIO control at individual RAN, while a centralised critic exploits graph-structured representations of inter-cell interactions during training. Graph attention mechanisms are employed to capture the heterogeneous and time-varying influence of neighbouring cells, improving learning stability and scalability. Simulation results demonstrate that the proposed approach achieves better performance than existing baselines.
Jinsheng Yuan, Mengbang Zou, Weisi Guo· International Mediterranean...· 0 citations