Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A deep reinforcement learning (DRL)-based adaptive routing scheme for maximizing throughput and minimizing end-to-end delay jointly in SAGIN and indicates that adaptive policy learning enables better congestion avoidance and more efficient resource utilization.
A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.
Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang· IEEE Access· 0 citations
A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.
M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al.· 0 citations
This paper proposes a Service Level Agreement (SLA)-aware resource allocation framework for 6G V2X slicing, realized as a Soft Actor-Critic (SAC) based xApp within the Open-Radio Access Network (O-RAN) near-realtime-RAN Intelligent Controller (near-RT-RIC). The xApp dynamically distributes radio resources across heterogeneous slices, minimizing SLA violations while considering fairness and throughput efficiency. Unlike heuristic or single-metric Deep Reinforcement Learning (DRL) methods, our design incorporates deadline awareness and service reliability directly into the reward formulation. Simulation results show that the proposed scheme consistently outperforms fixed, random, proportional, and Exponential moving Average (EMA)-based baselines, improving average packet delivery ratio (PDR), reducing mean SLA violations, and achieving a Pareto-optimal trade-off between throughput and compliance. These findings demonstrate the potential of O-RAN-native intelligent control for future 6G networks.
M. Tariq, Deepak Singh, M. Saad et al.· International Conference on...· 0 citations
Efficient coexistence of eMBB and URLLC services remains a critical challenge in AI-native Radio Access Networks (RANs). This paper proposes a two-timescale Hierarchical Reward Weighting (HRW) framework based on multiobjective reinforcement learning for context-aware O-RAN slicing under a Constrained Markov Decision Process (CMDP) formulation. The proposed architecture separates long-term policy adaptation from fast-timescale radio scheduling, mitigating the non-stationarity inherent in multiobjective RAN optimization. At the slow layer, a non-realtime RIC rApp exploits a long-term network context and a differentiable Softmax mapping to adapt slice reward preferences. These policies are propagated through the $O$ -RAN control hierarchy to guide downstream scheduling decisions. At the fast layer, decentralized scheduling agents embedded within the Open Distributed Unit (O-DU) MAC layer execute sub-millisecond Physical Resource Block (PRB) allocation and packet preemption, avoiding near-RT RIC transport latency constraints. Evaluated under a multiuser MIMO-OFDMA environment, the proposed framework improves resource utilization by up to 60.8% over static partitioning while maintaining bounded URLLC tail-latency behavior and strict Service Level Agreement (SLA) compliance. The results demonstrate the feasibility of AI-native hierarchical O-RAN control and align with the ITU-T visions for autonomous 6G RAN intelligence.
Charles Ssengonzi, Okuthe P. Kogeda, T. Olwal· 2026 ITU Kaleidoscope - AI a...· 0 citations
The experiments show that feasibility-aware learning can approach deterministic baseline reliability while retaining learned forwarding capability under hop constraints, and confirm that action masking is the dominant mechanism for maintaining feasible routing decisions, whereas trust mainly provides reliability-aware regularization.
Adeel Iqbal, Muhammad Faisal Siddiqui· Computers, Materials & C...· 0 citations