Jun 2026· IEEE Conference on Network Softwarization· pp. 180-188· 0 citations· 23 references
Computer Science
Abstract
G mobile networks are increasingly using Artificial Intelligence to manage highly dynamic environments characterized by time-varying traffic demands, user mobility, and heterogeneous resources. The dynamic behavior of User Equipment makes timely and accurate control decisions challenging, while distributed data exchange introduces communication overhead and privacy concerns. These challenges call for scalable and communication-efficient learning mechanisms for Radio Access Network (RAN) orchestration. In this paper, we propose DERRIC-FRL, a decentralized Federated Reinforcement Learning framework to orchestrate RAN intelligent controllers. DERRIC-FRL jointly optimizes controller placement and user power allocation through a selective two-level aggregation mechanism, reducing data exchange to only selected orchestrators and controllers while preserving user privacy and improving overall network performance. Specifically, our method significantly reduces total training communication costs by 34% across inter-domain connections, and up to 77% across intra-domain connections, compared to the FedAvg approach. Furthermore, DERRIC-FRL improves user throughput by up to 53% and 61% compared to the DERRIC and FedAvg baselines across a broad range of simulated scenarios.
Dynamic spectrum access (DSA) in 5G IoT setups with cognitive radio is characterized by rapid and decentralized decision-making processes in highly non-stationary wireless environments, limited communication needs, and restrictive bounds. In this work, we present F-DMRL, a federated, communication-efficient decentralized meta-reinforcement learning framework for allowing a massive number of IoT devices to meta-learn collectively about spectrum-access strategies in a decentralized way without centralized control and without an extensive amount of inter-agent communication. Our method incorporates lightweight federated meta-parameter aggregation with gradient sparsification and periodic communication, allowing devices to only compress the meta-updates during this process and then adapt locally for task-specificity. We have presented analytical speedup guarantees and upper bounds on communication cost under bounded environmental drift and shown that using the approach proposed here, F-DMRL preserves convergence properties while posing a large reduction in coordination overhead at the same time. Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation (up to 45% fewer episodes), higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines. Simulation results averaged across 10 independent runs demonstrate improvements of 45% faster adaptation and 60–80% lower communication overhead relative to baseline methods, while maintaining stable convergence.
Jayesh Kumar Dabi, Priyadarshi Ashok Dahat· International Journal of Wir...· 0 citations
X-CODE is an explainable offline MARL that operates offline without environmental interaction, nor inter-agent communication, nor inter-agent communication, and exploits explainability-aware reward shaping to modify the relative preference among joint offline transitions during centralized training to improve decentralized resource-allocation behavior.
Managing dynamic User Association and Resource Allocation (UARA) in modern Heterogeneous Cellular Networks (HetNets) remains a critical open challenge. Existing mathematical optimization and Reinforcement Learning approaches face limitations in handling low-latency decision-making under dynamic traffic conditions. This paper introduces a novel orchestration scheme for game-theoretic UARA in HetNets. The proposed bilevel framework distributes UARA decisions to User Equipment through a multi-objective non-cooperative game. Overlaying the distributed game, a centralized Deep Reinforcement Learning controller orchestrates network performance by dynamically configuring the game's utility parameters, enabling transitions between power awareness, coverage enhancement, and balanced operation. Evaluated on urban HetNet topologies with 3GPP TR 38.901-compliant channel modeling, the proposed framework closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods. Furthermore, it incurs low computational overhead and maintains stable performance across the evaluated traffic densities without retraining.
Sotiris Kopsinos, Alexandros I. Papadopoulos, Antonios Lalas et al.· 0 citations
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations
Results support the central conclusion that lightweight joint scheduling can materially improve wall-clock FL efficiency in heterogeneous 5G/6G edge networks.
Future 6G services will require strict performance guarantees, especially in terms of delay, end-to-end (e2e) across multiple network domains including packet and radio segments. While deterministic transport and slice-based capacity allocation can improve segment-level performance, ensuring e2e Network Service (NS) performance remains challenging as it requires making decisions Near–Real-Time (Near-RT) on a per-service basis, which does not fit well within the typical centralized control and orchestration hierarchy. Multi-agent systems (MAS), where a number of distributed agents collaborate, has demonstrated its capabilities for such Near-RT control. Agents equipped with Deep Reinforcement Learning (DRL) engines autonomously made traffic routing decisions based on e2e telemetry measurements. In this paper, we extend such MAS solutions for NS traffic routing focused on covering several issues that appear under frequent NS reconfiguration, e.g., caused by end device mobility. In addition, we define a lifecycle for NS operation that includes the initial MAS deployment, model reconfiguration during operation, and NS reconfiguration. The proposed lifecycle requires the definition of DRL training and validation procedures to produce models ready to be deployed with guaranteed performance under certain network conditions. In addition, model selection algorithms are defined for the lifecycle scenarios. In case of NS reconfiguration, a procedure for probe testing the actual network conditions is proposed to improve model selection. Evaluation across a meaningful set of network and traffic scenarios shows that the MAS is able to maintain e2e delay guarantees under all the lifecycle scenarios.
H. Shakespear-Miles, S. Barzegar, M. Ruiz et al.· IEEE Transactions on Network...· 0 citations