Jul 2026· International Conference on Edge Computing [Services Society]· pp. 226-234· 0 citations· 34 references
Abstract
The integration of 5G/6G networks with the Internet of Vehicles (IoV) requires efficient computational offloading for data-intensive applications such as autonomous driving and augmented reality. Although Unmanned Aerial Vehicles (UAVs) offer agile mobile edge computing (MEC) capabilities, their operational efficiency is hampered by high mobility, limited battery life, and the complexity of joint resource optimization. Existing offloading strategies often fail to simultaneously optimize latency, energy consumption, and resource utilization under dynamic IoV conditions. This paper proposes a novel Energy-Optimized Lightweight Deep Reinforcement Learning (DRL) framework for intelligent task offloading in UAV-assisted IoV networks. Our approach leverages a simplified Double Deep Q-Network (DDQN) to dynamically manage task partitioning by intelligent offloading decisions, UAV trajectory planning through optimized path forecasting, and resource allocation through adaptive computation distribution. Key innovations include a streamlined state-space design that reduces computational overhead by 30% and a composite reward function that balances latency and energy objectives. These are realized by a prioritized experience replay mechanism and a target network separation strategy that enhances learning stability. Experimental results demonstrate that our framework achieves a task success rate of 98.5%, reduces latency by 40%, and maintains a 78.1%. The results confirm the framework’s superiority, demonstrating significant improvements over its base architecture (DQN), its enhanced variant (DDQN), and other state-of-the-art baselines like MADDPG and game-theoretic approaches, thereby providing a robust solution for practical UAV-IoV deployments.
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.
Yongbin Wang, Peng Lin, Yan Liu et al.· IEEE Wireless Communications...· 0 citations
With the advancement of autonomous driving and smart navigation, Internet of Vehicles (IoV) systems face stringent requirements for real-time data delivery and processing reliability. Traditional metrics cannot fully capture information timeliness due to network dynamics and packet loss. Existing approaches also struggle with the coupling between task offloading and resource allocation, lacking adaptability in dynamic IoV environments. To address these issues, we propose a joint optimization scheme using a deep Q-network (DQN). Specifically, we build an IoV system model incorporating V2V and V2I communication, and formulate an optimization problem to minimize the average age of information (AAoI) under delay, bandwidth, computing, and energy constraints. We then design a mixed-action DQN algorithm with dual-network architecture, experience replay, and an action mask mechanism to enhance training stability and environmental adaptability. Simulation results show that our DQN-based scheme achieves the lowest AAoI among Random, Greedy, A2C, and DDQN, with reductions of 29.5%, 8.9 %, 7.1 %, and $\mathbf{7. 6 \%}$, respectively. It also exhibits superior delay and energy performance, confirming its effectiveness for dynamic IoV task offloading and resource allocation.
Chao He, Wanting Wang, Dongfeng Fu et al.· 2026 International Conferenc...· 0 citations
A new paradigm for satisfying the ever-growing demands of real-time Sixth Generation (6G) applications is Mobile Edge Computing (MEC). Additionally, base stations and Internet of Things devices that incorporate renewable energy harvesting capabilities have the potential to lower grid energy use. To maximize system potential and lower carbon emissions, it is crucial to make effective decisions about job offloading and resource allocation. A carbon-aware MEC architecture that uses both grid and renewable energy sources is proposed in this paper. Our goal is to jointly manage resource allocation and task offloading while monitoring carbon emissions and task queue delays to optimize system behavior under uncertainty, specifically for stochastic workloads and variable renewable generation. To balance these two cost components (emissions and queue length), we create a combined optimization problem. We develop a deep deterministic policy gradient (DDPG)-based joint optimization technique to address this issue in a constantly changing environment. In the optimization, we consider greedy policy (GP) and full offloading (FO), as well as time-average carbon emission (TACE) and time-average queue length (TAQL) as performance metrics, and time-average queue length (TAQL) and full execution (FE) as baseline strategies; we also evaluate normalized time-average cumulative reward (NTACR). This method uses continuous-action reinforcement learning to generate efficient, real-time control policies. For the proposed MEC network, numerical statistics show that our approach can lead to effective offloading and lower carbon emissions.
M. Saeed, Rashid A Saeed, M. A. Ahmed et al.· 2026 6th International Confe...· 0 citations
: The emergence of Unmanned Aerial Vehicle (UAV)-enabled Wireless Energy Transfer (WET) and Simultaneous Wireless Information and Power Transfer (SWIPT) technology provide a promising solution to overcome the energy sustainability limitations of traditional harvesting-reliant sensor networks. However, in large-scale Battery-free SWIPT-enabled Sensor Networks (BSSN) characterized by sparse node distribution and heterogeneous energy consumption and harvesting rates, employing a single UAV for energy replenishment often suffers from insufficient operation continuity and low charging efficiency. To overcome these challenges, a Multi-UAV Collaborative Energy Charging for BSSN Based on Multi-Agent Deep Deterministic Policy Gradient (MCEC-MADDPG) is proposed in this paper. Specifically, we construct a collaborative one-to-one precision energy supply model where UAVs hover directly above specific nodes to achieve power transmission without complex beamforming requirements. To achieve collaborative scheduling among multiple UAVs in wide-area dynamic environments, the energy replenishment problem is first formulated as a Partially Observable Markov Decision Process (POMDP). Subsequently, the Centralized Training with Decentralized Execution (CTDE) architecture of the MADDPG algorithm is leveraged to solve this POMDP, which effectively tackles the non-stationarity challenge inherent in multi-agent environments. Simulation results demonstrate that MCEC-MADDPG exhibits superior performance in terms of convergence speed and stability. It enables the adaptive emergence of spatial-division collaborative strategies, significantly enhances the average residual energy of the network, and elevates the node survival rate to nearly 90%. Compared with Deep Deterministic Policy Gradient (DDPG), the traditional static Partition-Greedy method, the heuristic K-Means algorithm and the dynamic Two-Layer task allocation strategy, the proposed approach demonstrates substantial advantages.
Xiangyi Le, Deyu Lin, Yufei Zhao et al.· Computers, Materials & C...· 0 citations