Skip to content
Open access

Enhanced Soft Actor–Critic with Dual-Path Channel Attention for UAV Autonomous Navigation in Complex Environments

Jul 2026 · Applied Sciences · Vol 16, pp. 7159 · 1 citation · 32 references

TL;DR

Compared with twin delayed deep deterministic policy gradient (TD3), SAC, and representative reinforcement learning methods such as attention-mechanism SAC based on attention mechanism enhancement, the proposed DPCA-ARF-PER-SAC method shows more balanced performance in task completion, navigation safety, and path efficiency.

Abstract

In complex and unknown environments, unmanned aerial vehicle (UAV) autonomous navigation still faces issues such as insufficient representation of state characteristics, fixed reward guidance, and low efficiency in utilizing key experience samples. To address these problems, this paper proposes an improved soft actor–critic (SAC) method that integrates dual-path channel attention (DPCA), adaptive reward feedback (ARF), and prioritized experience replay (PER); this method is named DPCA-ARF-PER-SAC. The proposed DPCA module is introduced into the actor network to recalibrate one-dimensional navigation state features and enhance the representation ability of key decision-making information. At the same time, the ARF mechanism can dynamically adjust the reward weights according to the training progress, while PER is used to improve the utilization efficiency of key samples. The experiments are conducted in a two-stage structure, including module-level ablation verification in the two-dimensional (2D) SimpleAvoid scenario and main performance comparison in the three-dimensional (3D) NH_center scenario. The experimental results show that the success rate of this method reaches 1.00 in the 2D scenario, and the collision rate is 0.00. In the 3D scenario, the success rate is 0.76, the collision rate is 0.24, and the average episode length is 232.8 steps. Compared with the baseline SAC, the success rate is increased by 2 percentage points, the collision rate is reduced by 2 percentage points, and the average episode length is reduced by 4.3 steps. Compared with twin delayed deep deterministic policy gradient (TD3), SAC, and representative reinforcement learning methods such as attention-mechanism SAC (AM-SAC) based on attention mechanism enhancement, the proposed DPCA-ARF-PER-SAC method shows more balanced performance in task completion, navigation safety, and path efficiency. These results indicate that DPCA-ARF-PER-SAC provides a more robust navigation strategy for complex 3D UAV autonomous navigation tasks.

Read PDF

Similar papers

Conference Aug 2026

Attention-based MADDPG with dual-buffer experience replay for cooperative multi-UAV target tracking

An Attention-based Multi-Agent Deep Deterministic Policy Gradient algorithm was developed for cooperative multi-unmanned aerial vehicle target tracking in dynamic environments and showed that the proposed algorithm improved convergence behavior and average reward, reduced the collision rate, and maintained competitive...

Qing-Lin Han, Hongmei Wang · 0 citations
Conference Aug 2026

A Multimodal Deep Reinforcement Learning Framework for Autonomous UAV Navigation in Gazebo–ROS Environments

Unmanned Aerial Vehicle (UAV) autonomous navigation is a key capability for UAVs, particularly when in complex and dynamic environments where continuous human control is impractical. While commonly used rule-based navigation and path-planning techniques can be effective in structured environments, they can be ineffecti...

B. S. Santhoshi, B. S, R. Shankar · 0 citations
Sep 2026

A Multi-UAV Cooperative Navigation Method Based on Policy Decomposition Structure

Cooperative navigation of multiple unmanned aerial vehicles (UAVs) in disaster search-and-rescue scenarios is challenging due to dense obstacles, partial observability, and strong inter-agent coupling, which often result in path conflicts, collision risks, and limited policy generalization. To address these challenges,...

Li Tan, Hai-Xia Zhao, Jia-Qin Chai et al. · 0 citations
Sep 2026

AUVs Path Planning Based on DSAC-T With Reward-Adaptive PER

Efficient 3-D path planning for autonomous underwater vehicles (AUVs) in dynamic submarine environments presents a significant challenge due to complex seabed terrain, ocean currents, and obstacles. In view of the adaptability and generalization limitations of traditional methods, this article proposes the reward-adapt...

Xin Cheng, Hai Jin, Yun Chen et al. · 0 citations
Open access Aug 2026

Modeling Dynamic Obstacle Avoidance Strategy of Drone Swarms Combined with Multi-Agent Reinforcement Learning

The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.

X.-H. Fang, K. Chen, Cheng-Hao Ren et al. · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.