Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 18 references
Abstract
Robust Reinforcement Learning (RL) based task scheduling approaches can address the inherent tradeoff between energy consumption and deadline violation in a Multi-access Edge Computing (MEC) based Internet of Things (IoT) network, while maintaining robustness against changes in the task arrival rate. However, tabular robust RL algorithms suffer from high computational and storage complexity, and therefore are not scalable to a system with a large number of IoT nodes. To this end, in this paper, we propose a robust deep RL based task scheduling algorithm to solve the underlying Robust-Return Constrained Markov Decision Process (R2CMDP) problem. The proposed algorithm introduces a tunable amount of robustness in the solution of the RL framework. Complexity analysis and ns-3 simulation results are presented to demonstrate the efficacy of our algorithm.
Results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.
L. Hoang, Van-Tam Hoang, Huu-Huy Ngo· International journal of Com...· 1 citation
A dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs and achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods.
Suresh Betam, S. Nagendram, Bathula Prasanna Kumar et al.· Scientific Reports· 0 citations
The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.
Oussama Lagnfdi, Marouane Myyara, A. Darif· International journal of Com...· 0 citations
This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.
Olumide Alamu, T. Olwal, Emmanuel M. Migabo· Network· 0 citations
Comparative tests with PPO, FIFO, FAIR and HAS baselines confirm that multi-agent reinforcement learning can well capture the intrinsic scheduling patterns of complex mobile environments, providing an adaptive and energy-efficient scheduling solution for practical IoT deployments.
Haoyu Gu· Scientific Journal of Intell...· 0 citations
An intelligent routing algorithm called Reinforcement Learning-based Congestion-Aware Routing (RLbCAR) is introduced for intelligent routing in IoT sensor networks and ensures reliable, congestion-adaptive, and computationally efficient routing in a resource-limited IoT sensor network.
M. Sunitha, M. Prashanth, Yenugula Swapna et al.· Discover Computing· 0 citations