Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 56 references
TL;DR
A two-level Q-learning-based geographic routing protocol called TLQ-Geo for FANETs, which significantly reduces convergence time and computational overhead and integrates hierarchical decision-making with adaptive reinforcement learning.
Abstract
In flying ad hoc networks (FANETs), high node mobility, dynamic topology, and limited resources, such as energy and bandwidth, lead to unstable links and short-lived routes. In such an environment, although Q-learning-based routing methods are adaptable, they face serious challenges in practice due to large state space, high computational load, and slow convergence. To address these issues, this paper proposes a two-level Q-learning-based geographic routing protocol called TLQ-Geo for FANETs. This protocol integrates hierarchical decision-making with adaptive reinforcement learning. TLQ-Geo divides the routing process into two layers: the guided region selection (GRS) layer and the Q-learning-based routing (QRL) layer. The GRS layer determines a bounded search corridor between the source and the destination using a chain of intelligent decision points (IDPs), while the QRL layer performs distributed path optimization within this virtual corridor via Q-learning. By restricting the state space to the region guided by IDPs, TLQ-Geo significantly reduces convergence time and computational overhead. In addition, a dynamic inter-layer feedback mechanism periodically evaluates the performance of each IDP chain and adaptively reconfigures it under topology variations. Extensive simulations demonstrate that when the node density varies, TLQ-Geo achieves higher network lifespan (approximately 4.51%), improved packet delivery ratio (about 1.25%), lower routing overhead (around 3.38%), and better energy efficiency (about 20.79%), while the delay increases by about 13.84%, compared to three basic routing methods, namely QRCF, QRF, and QFAN. Also, when the node speed changes, TLQ-Geo yields better network lifespan (approximately 5.46%), higher packet delivery ratio (about 1.69%), lower overhead (around 2.80%), and better energy efficiency (about 6.42%), while the delay increases by about 9.09%.
The paper introduces RML-ZEREM to solve existing limitations, which functions as a Reinforcement Learning (RL) based Zone-Based Leader-Aware Energy-Efficient Routing Protocol for MANETs, which serves next-generation MANET applications.
Rani Sahu, Babita Rathore· Journal of Intelligent Compu...· 0 citations
Simulation results obtained demonstrate that Q-WeCBR outperforms CBR, DSDV, and GPSR in terms of packet delivery ratio and throughput, confirming the effectiveness of clustering combined with learning-based routing for dynamic vehicular networks.
Ahlam Boussadia· International journal of inf...· 0 citations
Mobile Ad Hoc Networks (MANETs) are expected to support highly dynamic and decentralized communication scenarios in future 6G-oriented wireless systems. However, routing remains challenging because of mobility, topology variability, and resource constraints. Reinforcement learning (RL) offers a promising alternative by enabling adaptive routing decisions based on observed network conditions. This paper presents QL-6GRP (Q-Learning for 6G Routing Protocol), a lightweight Q-learning-based routing protocol designed for fully distributed MANET environments. The protocol enables each node to learn next-hop forwarding decisions using local observations, including link quality, residual energy, hop progress, and neighborhood density. A complete implementation of QL-6GRP was developed within the NS-3 simulator, supporting online learning through hop-level feedback signaling and bounded-memory operation. The protocol was evaluated under multiple parameter settings and network sizes using Random Waypoint mobility and UDP constant-bit-rate traffic to examine both routing performance and computational behavior. The experimental results demonstrate the feasibility of adaptive routing with moderate signaling overhead under carefully tuned moderate-scale scenarios while revealing key trade-offs between feedback frequency, routing quality, and computational scalability. Moderate periodic feedback provides the most favorable balance, whereas excessive feedback increases overhead without improving performance. In addition, reinforcement-learning operations incur substantial computational costs, with the wall-clock runtime increasing by approximately 13 times when the network size increases from 50 to 100 nodes. These findings reveal the operating limits and practical design trade-offs of lightweight tabular RL-based MANET routing and provide useful guidelines for future scalable learning-driven protocols in dynamic wireless environments.
Simulation results indicate that HOA-MEPFL-CLCT-RP outperforms existing models in terms of Packet Delivery Ratio (PDR), energy efficiency, End-to-End Delay (E2D), and routing overhead.
Shaleena H, Sumangala K· International journal of com...· 0 citations
Multipath routing in wireless sensor networks (WSNs) improves reliability by providing alternative forwarding paths when a route fails. However, mobile sinks make path maintenance difficult because sink movement can invalidate previously constructed source-to-sink routes. Existing protocols typically depend on either global path reconstruction, which increases control overhead, or footprint-chaining, which accumulates detours through previous sink positions and may weaken path independence. To address this problem, this paper proposes QL-LGMPRP, a reliability-aware local-grid-based multipath routing protocol that combines a sink-centered local grid, two-path delivery, link-quality-aware forwarding, and lightweight tabular Q-learning for waypoint adaptation. Mobility-related route changes are confined to the sink-centered grid, whereas a compact tabular Q-learning policy adjusts the primary-path direction using grid, link-quality, and energy-related state variables. The sink constructs a local grid around its current position, with cells sized to keep in-grid forwarding locally bounded. When an event occurs, the source computes an entry point on the grid perimeter and constructs two greedy paths: a primary path through a Q-learning-selected waypoint near the grid boundary and a backup path toward the current sink position. The Q-learning agent uses a compact tabular state representation that includes the boundary-cell index, residual-energy level, sink-grid position, and local link-quality information, and learns waypoint offsets using a reward that combines delivery success, transmission energy, and delay. This design confines routing adaptation to the sink-centered grid while allowing the waypoint policy to respond to heterogeneous link conditions. Simulation results under different sink speeds and interference conditions show that QL-LGMPRP maintains high delivery reliability while reducing detour-related forwarding costs relative to footprint-chaining and showing lower weak-link exposure than the geometric-forwarding comparison schemes.
Cheonyong Kim, Sangdae Kim· Applied Sciences· 0 citations