Aug 2026· Journal of universal computer science (Online)· 0 citations· 30 references
TL;DR
This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.
Abstract
To address the growing need for wireless communications energy efficiency, this paper proposes a new multi-agent reinforcement learning (MARL) approach to cooperative data offloading in heterogeneous cellular networks. This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions. We formulate the offloading problem as a Markov Decision Process (MDP) with correct models of energy consumption and network conditions. The deep Q-network (DQN)-based MARL algorithm allows user equipment (UEs) to learn collaborative strategies for optimizing overall energy consumption and timely offloading of data. Simulations compare MARL against greedy, random, and independent Q-learning baselines in low and high mobility regimes. Experiments show that MARL saves energy by as much as 40% over random offloading and 16.6% over greedy offloading, as well as enhancing average delay, throughput, and fairness. Convergence of the learning rate of the algorithm is within 1000 episodes, and sensitivity analyses confirm its performance across a range of user density and data size settings. Furthermore, the MARL framework accommodates dynamic network conditions and provides an adaptable solution for network operators to maximize performance and sustainability for existing and future wireless networks. The suggested MARL framework still performs better in terms of delay, throughput, and fairness, while at the same time catapulting energy savings over the greedy offloading approach by 9.4% to 12.8% under the very realistic 3GPP-compliant HARQ, adaptive modulation, standardized power control, and urban SLAW mobility models settings. With the extended state/action spaces and realistic 3GPP+SLAW conditions, MARL has an energy savings of 11.7%–14.2% over greedy offloading while still having the leading delay, throughput, and fairness.
This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.
Mugerwa Joseph, Ajaegbu Chigozirim· International Journal Of Eng...· 0 citations
Simulation results position D3QN-PER as a strong candidate for deployment as a near-RT RIC xApp within the O-RAN architecture, advancing the vision of AI-native mobility management for 6G.
Kalpesh Popat, Divyakant T. Meva· Telecommunications Systems· 0 citations
Cognitive radio networks (CRNs) has significant potential for optimizing spectrum usage, but they still face issues, such as high energy consumption, increased transmission delay, and inefficient routing due to unpredictable network performance and spectrum variance. Relatively new developments, including deep reinforcement learning (DRL), allow for simultaneous routing and resource management, but still need substantial number of trial and error interactions with the environment, which can consume energy and lead to a significant convergence time. Maintaining optimal connectivity with both primary users (PUs) and secondary users (SUs) in large scale CRNs following a homogenous Poisson process while optimally utilizing spectrum access remains a challenge. Towards advancing these problems, we propose an energy-aware, cross-layer routing design based on an apprenticeship learning framework. Our multi-stage dynamic adjustment rating (DAR) mechanism allows for effective tuning of transmit power to decrease action space (following a multi-level transition) and reduce energy consumption. The Whale Swarm Optimization Algorithm (WsOA) allows us to provide predictions of link connectivity and path probability to ensure a more reliable routing selection is presented. To ensure faster convergence and a minimal memory footprint, we incorporate Deep Q-learning with a rapidly converging local search method based on permutation-equivariant neural networks for unseen environments in the given network scenarios. The simulation results show that the suggested approach outperforms conventional algorithms like CRQ-routing, PM-DQN, and other DQL methods in terms of throughput (100-70) pt/s, routing delay (5-1) ms, packet delivery ratio (100-70) % and Percentage packet loss (5-1) %.
M. Lakshmi, Arram. Mahesh Babu· Journal of Circuits, Systems...· 0 citations
Vehicular edge computing (VEC), a key enabler for the Internet of Things (IoT) in intelligent transportation, addresses onboard processing constraints through collaborative task offloading among vehicles, facilitating latency-sensitive applications such as autonomous driving. However, developing efficient offloading strategies remains particularly challenging in high-density vehicular networks, where intensive computational demands coexist with severely constrained intervehicle communication ranges due to signal blockage. To handle this, we propose M4O, a mobility-aware task offloading framework supporting multihop, multiuser, and multitask offloading optimization. M4O intelligently integrates vehicle mobility patterns and enables relay-assisted offloading to enhance system effectiveness and robustness. The framework employs a dual-algorithm approach: the advantage actor–critic (A2C) for indivisible tasks and the hybrid proximal policy optimization (H-PPO) for divisible tasks, both optimized to minimize the temporally coupled composite cost of time and resources. Extensive experiments demonstrate that the deep reinforcement learning (DRL)-based solutions of M4O deliver stable and efficient offloading strategies, outperforming existing benchmarks by significant margins in cost efficiency. Our code is available at https://github.com/Zhouym1028/M4O
Momiao Zhou, Yimin Zhou, Yanshi Sun et al.· IEEE Internet of Things Jour...· 0 citations
Integrated sensing, communication, and computation (ISCC) enables next-generation wireless networks to perform environmental perception while processing massive data under stringent quality-of-service (QoS) requirements. Energy consumption is a crucial indicator for the ISCC system design. However, accounting for energy heterogeneity in ISCC system design is an open problem. Specifically, battery-constrained user equipments (UEs) and energy-abundant access points (APs) require fundamentally different energy allocation strategies based on device computational capabilities, battery states, and QoS constraints. In this paper, we introduce a nonconvex energy cost minimization problem by considering a user-specific energy cost ratio coefficient that explicitly balances UE-AP energy consumption according to heterogeneous device energy states. To efficiently address this problem, a double-loop framework combining successive convex approximation and alternating direction method of multipliers is also developed. Numerical results demonstrate that the proposed scheme significantly outperforms the fixed offloading baselines (full offloading, full local and half offloading) in terms of the total energy cost. In particular, the proposed scheme achieves up to $25-47.6\%$ energy cost reduction at moderate latency constraints over fixed offloading baselines, thereby supporting time-sensitive applications. Moreover, this work provides an effective solution for energy-efficient and QoS-aware 6G ISCC systems serving diverse devices with conflicting energy priorities.
Kai Dong, Lei Wang, S. Vorobyov et al.· IEEE Transactions on Wireles...· 0 citations
This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.
Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al.· International journal of Com...· 0 citations