Adaptive Action-Constraint Safe Driving Decision Control Algorithm Based on Deep Reinforcement Learning
Abstract
Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods are developed for standardized domains such as games, with limited work tailored to the highly dynamic and safety-critical context of autonomous driving. In this paper, an enhanced deep reinforcement learning algorithm is proposed based on the Proximal Policy Optimization (PPO) framework, incorporating trajectory-based adaptive clipping and expert guidance to address the challenges of complex autonomous-driving environments. An adaptive clipping threshold is constructed using historical trajectory information, effectively reducing policy oscillations and improving both the stability and convergence speed of learning in dynamic scenarios. In addition, imitation learning (IL) is incorporated into the reinforcement learning process, enabling expert demonstrations to guide exploration and thereby improving sample efficiency. The effectiveness of the proposed algorithm was evaluated on six maps in the CARLA simulator and compared against the vanilla PPO baseline. The results show that the improved algorithm achieves faster training, better performance, and produces a smoother policy.