Results highlight the effectiveness and practicality of the proposed HybridRL-RNP framework as an intelligent topology control solution for Wireless Mesh Networks towards 6G, where the synergy between AI-driven optimization and heuristic knowledge plays a pivotal role in achieving globally optimal network connectivity.
Abstract
In Wireless Mesh Networks (WMNs), router node placement (RNP) is critical for achieving network-wide coverage, connectivity, and reliable communication performance. However, determining the optimal router placement is a nonlinear combinatorial problem with a vast and dynamic search space, which is strongly influenced by user distribution and traffic demand. To address this challenge, this study introduces a Hybrid Reinforcement Learning Framework, called HybridRL-RNP, designed to optimize router placement in WMNs. The proposed approach integrates the REINFORCE algorithm with a heuristic-guided placement strategy, enabling the learning agent to adaptively select router positions based on the current network coverage and the node density. A reward function is formulated using network connectivity (NC), which is the proportion of user nodes connected to at least one gateway as the primary optimization objective, while also considering coverage uniformity and router interconnectivity. Simulation experiments demonstrate that HybridRL-RNP achieves an average NC exceeding 96%, outperforming traditional heuristic-based and pure RL-based schemes. Moreover, the framework ensures stable inter-router topologies and scalable coverage performance for varying network densities. These results highlight the effectiveness and practicality of the proposed HybridRL-RNP framework as an intelligent topology control solution for Wireless Mesh Networks towards 6G, where the synergy between AI-driven optimization and heuristic knowledge plays a pivotal role in achieving globally optimal network connectivity.
The experiments show that feasibility-aware learning can approach deterministic baseline reliability while retaining learned forwarding capability under hop constraints, and confirm that action masking is the dominant mechanism for maintaining feasible routing decisions, whereas trust mainly provides reliability-aware regularization.
Adeel Iqbal, Muhammad Faisal Siddiqui· Computers, Materials & C...· 0 citations
This review analyses topology-aware learning-based routing for STINs, concentrating on Graph Neural Networks (GNNs) and hybrid GNN–Reinforcement Learning (GNN–RL) frameworks.
Eyeneka J. Ntuen, A. Obot, K. Udofia et al.· International journal of re...· 0 citations
A deep reinforcement learning (DRL)-based adaptive routing scheme for maximizing throughput and minimizing end-to-end delay jointly in SAGIN and indicates that adaptive policy learning enables better congestion avoidance and more efficient resource utilization.
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations
Networks with highly dynamic data transmission demands and network topologies are common in real world. A fundamental problem in such networks is achieving scalable traffic allocation to maximize long-term total throughput under link capacity constraints. However, state-of-the-art (SOTA) works lack scalability. This is primarily due to two reasons in large-scale networks: first, they require solving constrained optimization problems online, which leads to high decision latency; second, they rely on reinforcement learning algorithms for policy optimization, which are inefficient in exploration and challenging to train effectively. To address these issues, we propose the Fast Networked Control (FNC) policy framework, which firstly utilizes parallelizable neural network modules to process the state and generate raw decisions, followed by basic operations such as normalizations and comparisons, which do not require iteration or optimization, to obtain decisions that satisfy the constraints. Hence, FNC policy avoids solving constrained optimization problems and supports parallel execution, significantly reducing decision latency. Furthermore, this policy preserves gradient flow and supports backpropagation, which enable us to design an imitation learning algorithm to efficiently train the policy in an end-to-end manner. Experiments in large-scale networks show that our FNC policy achieves an average 8% improvement in demands satisfaction and 10 times reduction in decision latency versus SOTA works.
Zhaoxing Yang, Guiyun Fan, Anjie Cao et al.· IEEE Transactions on Network...· 0 citations
This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.
Berke Kilinç, M. Ö. Efe· International Conference on...· 0 citations