A hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks and achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.
Abstract
Existing studies on deep reinforcement learning based driving controllers often focus on traffic scenarios with relatively simple patterns. This limits their ability to handle challenging highway overtaking scenarios with delayed long-term rewards and reduces the generalizability of the learned policy. This paper presents a hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks. The framework contains a high-level controller for long-term planning and exploration and a low-level controller for detailed longitudinal and lateral control. To better expose the high-level controller to delayed rewards, the two controllers are trained separately in a two-step process. In the first step, the high-level controller is trained with a critic-gated goal completion mechanism and a fixed rule-based low-level motion planner. In the second step, the trained high-level policy guides the learning of the low-level controller. To evaluate long-horizon exploration, we design a speed-biased highway reward and a highway overtaking trap scenario involving both slower-moving vehicles and general traffic vehicles. We conduct the experiments in the highway-env simulation environment. Compared with Double DQN, the best-performing single-level controller design, our hierarchical framework improves trap-escape success from 0% to 100%, increases average speed from 10.63 m/s to 14.26 m/s, and increases the traveled distance from 248.17 m to 332.83 m during the fixed 50 decision steps long experimental episode window. Our method also achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.
Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods...
Rui Guo, Xin-Yu Li, Zhong-Hao Fu et al.· IEEE Open Journal of Intelli...· 0 citations
A hierarchical framework for behavior decision-making and motion planning that explicitly accounts for the dynamic balance between safety and driving efficiency is proposed, and achieves a highly comparable safety-efficiency trade-off frontier to constrained MDP benchmarks but under a single unified network framework.
Wei Liu, Yong-Qing Jia, Chu-Dong Lin et al.· Proceedings of the Instituti...· 0 citations
This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshe...
This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker with collaboration between a high-level manager and a low-level worker.
Qi-Ming Chen· International Conference on...· 0 citations
Planning Diffusion Policy Optimization is proposed, an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation and obtains an improved success rate over strong baselines and ablations demonstrate that action chunks are especially impor...
This work presents NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking that encodes full-environment obstacle and boundary geometry in a target-relative vertex representation, allowing the policy to retain route-defining context throughout the approach.
Zihan Wang, Baixiang Huang, Yang Guan et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.