A dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs and achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods.
Abstract
Internet of Things-based wireless sensor networks (IoT-WSNs) face persistent challenges related to energy consumption, latency, and network congestion under dynamic and heterogeneous topologies. Conventional reinforcement learning approaches rely on static reward formulations, which limit adaptability and hinder effective multi-objective optimization. This study proposes a dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs. The proposed approach employs real-time reward recalibration to jointly optimize energy efficiency, delay, and throughput under varying network conditions. A hybrid deep reinforcement learning architecture is developed by integrating value-based, policy-based, and actor-critic methods, along with multi-agent coordination and attention mechanisms to prioritize critical nodes and links. Furthermore, a hierarchical learning structure decomposes global and local routing objectives, improving scalability and decision efficiency in complex network environments. Experimental results demonstrate that the proposed framework achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods. These findings highlight the effectiveness of dynamic reward adaptation for scalable and robust multi-objective optimization in IoT-WSN routing.
Wireless Sensor Networks (WSNs) play a crucial role in the expanding landscape of the Internet of Things (IoT), yet they continue to face persistent challenges related to energy consumption, computational efficiency, and scalability. Although protocols like the Energy-Efficient Routing Protocol through Hybrid Algorithms (EERHA) have made progress in extending network lifespan using machine learning, they still encounter limitations particularly in managing processing overhead, adapting to changing conditions, and scaling to larger deployments. Unlike existing approaches that merely combine machine learning modules, DEERL-WSN (Distributed Energy-Efficient Reinforcement Learning for Wireless Sensor Networks), introduces a unified hierarchical learning framework where distributed reinforcement learning, lightweight graph neural networks, transfer learning, and federated optimization operate cooperatively. The novelty lies in (i) adaptive reward-driven routing using meta-learned objective weights, (ii) topology-aware clustering through lightweight GNN embeddings with significantly reduced computational complexity, (iii) transfer learning-assisted cluster-head prediction to eliminate repetitive optimization overhead, and (iv) federated deep reinforcement learning enabling scalable learning without centralized processing bottlenecks. These integrated innovations collectively address energy efficiency, scalability, and computational constraints simultaneously, which remain largely unresolved in existing WSN routing frameworks. DEERL-WSN addresses several bottlenecks found in earlier protocols and the simulation results show that DEERL-WSN significantly outperforms EERHA and other state-of-the-art methods.
Maheshkumar Patil, B. J, K. R et al.· 2026 7th International Confe...· 0 citations
Energy remains the most critical and limiting resource in Wireless Sensor Networks (WSNs) and Internet of Things (IoT) systems, directly constraining network lifetime, scalability, and real-world deployability. Although multi-hop routing is widely adopted to reduce transmission energy and balance traffic load, recent solutions increasingly rely on metaheuristic optimization and machine learning techniques whose computational, control, and learning overhead is rarely accounted for. This leads to a fundamental energy–intelligence trade-off that challenges the sustainability of intelligent routing in resource-constrained environments. This paper presents a critical, energy-centric review of multi-hop routing approaches for IoT and WSNs proposed between 2018 and 2025. Heuristic, metaheuristic, dynamic and Heterogeneous routing, reinforcement learning, deep reinforcement learning, and explainable AI-based protocols are systematically analyzed with an emphasis on net energy efficiency, scalability, feasibility on constrained devices, and model realism, rather than reported performance gains alone. The analysis reveals that energy is predominantly treated as a secondary optimization objective rather than as a governing system constraint. To address this limitation, we propose a hybrid and explainable routing framework governed by energy awareness, in which intelligence activation is explicitly conditioned on its net energy benefit. This perspective provides a principled foundation for sustainable and trustworthy intelligent routing in next-generation IoT and WSN systems.
Moez Elarfaoui, Hamdi Ouechtati, Nadia Ben Azzouna· International Conference on...· 0 citations
Results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.
L. Hoang, Van-Tam Hoang, Huu-Huy Ngo· International journal of Com...· 1 citation
This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.
Olumide Alamu, T. Olwal, Emmanuel M. Migabo· Network· 0 citations