Skip to content
Open access

RIS assisted energy aware multi agent adaptive SAC for UAV aided IoT network path planning and obstacle avoidance.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

This study introduces a multi agent soft actor-critic (MASAC) framework for UAV path planning and energy aware coordination in a fixed RIS assisted IoT grid, and offers a simulation level benchmark for energy efficient UAV navigation in RIS assisted IoT environments.

Abstract

Unmanned aerial vehicles (UAVs) are emerging as critical enablers of next generation Internet of Things (IoT) infrastructures, supporting real-time data collection, wireless relaying, and agile operations in dynamic environments. However, achieving safe and energy efficient multi UAV navigation under sixth generation (6G) communication constraints remains a significant challenge due to dynamic obstacles, limited on board energy, and the high cost of centralized coordination. This study introduces a multi agent soft actor-critic (MASAC) framework for UAV path planning and energy aware coordination in a fixed RIS assisted IoT grid. MASAC integrates entropy regularized actor-critic learning with reconfigurable intelligent surface (RIS) aware reward shaping to support energy aware navigation, RIS assisted recharging, and connectivity guided trajectory optimization. A lightweight convolutional policy network is used to encode spatial information from the grid environment, including obstacle locations, dynamic obstacle states, exploration memory, and UAV position, enabling efficient policy learning under constrained navigation settings. Extensive simulations in RIS assisted, 6G enabled IoT environments demonstrate that MASAC achieves a 100% mission success rate, where mission success is defined as reaching the fixed goal cell before energy depletion and within the maximum episode horizon of 500 steps. Compared with the strongest baseline success rate of 75%, this corresponds to a 25%-point absolute improvement and a 33.3% relative improvement under the same evaluation protocol and identical environmental settings. Within the adopted grid level energy abstraction, MASAC also achieves approximately 33% higher RIS recharge utilization. It also provides 6% greater grid level 6G connectivity and 23% higher cumulative reward. Meanwhile, it maintains a low simulation time evaluation latency of approximately 38 ms per UAV. Statistical analysis confirms these gains as significant ([Formula: see text]). The proposed framework offers a simulation level benchmark for energy efficient UAV navigation in RIS assisted IoT environments. It also supports future deployment oriented research under realistic operational constraints.

Read PDF

Similar papers

Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
Open access 2026

A Novel Deep Reinforcement Learning Approach for UAV Path Planning and Obstacle Avoidance for IoT Data Collection in Complex Urban Environment

In uncrewed aerial vehicle (UAV)-assisted Internet of Things (IoT) networks, UAVs often need to fly at low altitudes and navigate through obstacles to maintain reliable communication with IoT nodes during data collection missions. This paper proposes a novel deep reinforcement learning (DRL)-based approach for 3D UAV path planning and obstacle avoidance, with the objective of minimizing data collection time from IoT nodes distributed across complex urban environments. To address this problem, we propose an improved DRL algorithm, Dropout-based Prioritized Soft Actor-Critic (DPSAC), which integrates the Soft Actor-Critic (SAC) algorithm with Prioritized Experience Replay (PER) and the Dropout technique. Furthermore, two innovative approaches are introduced to enhance the algorithm’s performance. First, the Episodic Environment (EN) training approach introduces random variations in obstacle number, position, and height across training episodes, thereby enhancing the agent’s ability to generalize its learned policy to new and unknown environments. Second, the Switching Reward mechanism reduces penalties for collisions and boundary violations in the reward function during the early stages of training, thereby facilitating exploration and accelerating the agent’s learning of IoT-related tasks. Simulation results demonstrate that the proposed DRL-based approach achieves faster convergence and greater stability during the training process compared to baseline algorithms. Specifically, experiments conducted in new and complex environments show that this method can collect data with an average collision-free success rate of 98% from 10 IoT nodes and 95% from 20 IoT nodes, confirming its remarkable superiority over the baseline algorithms.

M. Jenabi, Hadi Asharioun, M. Pourgholi · 0 citations
Open access Aug 2026

AI-Driven Energy-Efficient Routing and UAV Trajectory Optimization for UAV-Assisted Internet of Things Sensor Networks in 6G Environments

The results confirm that the integration of artificial intelligence, energy-aware routing, and UAV trajectory optimization provides an effective and scalable solution for next-generation UAV-assisted IoT systems and establishes a robust foundation for intelligent 6G-enabled wireless sensor networks.

Mojtaba Nasehi · 0 citations
2026

Energy-Aware Multi-UAV Collaboration for Data Collection and Trajectory Planning With MADDPG

Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.

Jing Mei, Jinglei Xu, Zhao Tong et al. · 0 citations
2026

Safety-Constrained UAV Trajectory Planning for AoI Minimization in Post-Disaster IoT Networks

Unmanned aerial vehicle (UAV)-assisted Internet of Things (IoT) data collection is a promising solution for timely information acquisition in post-disaster scenarios with damaged terrestrial infrastructure. However, freshness-aware UAV trajectory planning is challenging due to the coupled effects of heterogeneous ground node priorities, Age of Information (AoI) evolution, continuous UAV control, and safety risks caused by no-fly zones and initially unknown obstacles. In this letter, we formulate the safety-constrained weighted AoI minimization problem as a constrained Markov decision process (CMDP) and propose a safety-constrained twin delayed deep deterministic policy gradient (SC-TD3) algorithm with Lagrangian safety optimization to decouple the AoI-oriented objective from long-term safety-risk control and adaptively balance information freshness and safety risk during policy learning. Simulation results show that SC-TD3 achieves higher accumulated reward and reduces mean weighted AoI by 64.3%–78.2% and 67.1%–73.6% in the CN-ratio and GN-scale tests, respectively, while reducing mean total safety cost by 61.4%–75.6% compared with the strongest benchmark algorithm.

Jinghao Wang, Xu Wang, Jihao Luo et al. · 0 citations
2026

AoI-Aware UAV-Assisted Secure Status Updating: An Agentic AI-Enabled DRL Approach

The rapid expansion of real-time Internet of Things (IoT) applications has positioned uncrewed aerial vehicles (UAVs) as a promising solution for flexible and timely data collection in areas lacking robust infrastructure. This paper investigates a UAV-assisted secure status updating system, where a UAV serves as a mobile relay to forward status updating packets from ground devices (GDs) under the threat of a potential eavesdropper. To ensure information freshness and operational sustainability, we formulate a long-term stochastic optimization problem to minimize the cumulative average age-of-information (AoI) of all GDs and energy consumption of the UAV. The formulated optimization problem is an online mixed-integer non-linear programming problem, which involves the joint optimization of the flight speed, direction, and transmission power of the UAV as well as the binary scheduling indicator of GDs. To tackle the inherent non-convexity and complex spatial-temporal coupling, we propose an agentic artificial intelligence (AI)-enabled deep reinforcement learning (DRL) approach, named adaptive truncated quantile critics with large language models (LLM)-enabled state representation and reward function design (ATQC-L). Specifically, an adaptive truncated quantile mechanism is incorporated to mitigate distributional overestimation in dynamic environments. Furthermore, we leverage the reasoning capability of LLMs as an offline design-time agent to generate task-aware state representation and intrinsic reward functions. Simulation results demonstrate that the proposed ATQC-L algorithm outperforms representative DRL baselines in balancing information freshness and energy consumption of the UAV, while maintaining stable performance under different network scales, LLM backbones, truncation-parameter settings, imperfect eavesdropping channel state information, and mobile eavesdropping scenarios.

Chuang Zhang, Geng Sun, Jiahui Li et al. · 0 citations