Skip to content
Open access

Extensive Exploration in Highway Overtaking Scenarios Using Hierarchical Reinforcement Learning

2026 · IEEE Access · Vol 14, pp. 123965-123978 · 0 citations · 44 references
Computer Science

TL;DR

A hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks and achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.

Abstract

Existing studies on deep reinforcement learning based driving controllers often focus on traffic scenarios with relatively simple patterns. This limits their ability to handle challenging highway overtaking scenarios with delayed long-term rewards and reduces the generalizability of the learned policy. This paper presents a hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks. The framework contains a high-level controller for long-term planning and exploration and a low-level controller for detailed longitudinal and lateral control. To better expose the high-level controller to delayed rewards, the two controllers are trained separately in a two-step process. In the first step, the high-level controller is trained with a critic-gated goal completion mechanism and a fixed rule-based low-level motion planner. In the second step, the trained high-level policy guides the learning of the low-level controller. To evaluate long-horizon exploration, we design a speed-biased highway reward and a highway overtaking trap scenario involving both slower-moving vehicles and general traffic vehicles. We conduct the experiments in the highway-env simulation environment. Compared with Double DQN, the best-performing single-level controller design, our hierarchical framework improves trap-escape success from 0% to 100%, increases average speed from 10.63 m/s to 14.26 m/s, and increases the traveled distance from 248.17 m to 332.83 m during the fixed 50 decision steps long experimental episode window. Our method also achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.

Read PDF

Similar papers

Open access 2026

Adaptive Action-Constraint Safe Driving Decision Control Algorithm Based on Deep Reinforcement Learning

Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods...

Rui Guo, Xin-Yu Li, Zhong-Hao Fu et al. · 0 citations
Sep 2026

Behavior decision-making of intelligent vehicles using budgeted reinforcement learning

A hierarchical framework for behavior decision-making and motion planning that explicitly accounts for the dynamic balance between safety and driving efficiency is proposed, and achieves a highly comparable safety-efficiency trade-off frontier to constrained MDP benchmarks but under a single unified network framework.

Wei Liu, Yong-Qing Jia, Chu-Dong Lin et al. · 0 citations
Open access Aug 2026

Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation

This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshe...

Yi-Tan Wang, Yang-Yang Zhang, Zhen-Xing Gao · 0 citations
#reinforcement learning Conference Sep 2026

A hierarchical reinforcement learning approach for robot navigation integrating LiDAR priors and the options framework

This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker with collaboration between a high-level manager and a low-level worker.

Qi-Ming Chen · 0 citations
Preprint Aug 2026

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

Planning Diffusion Policy Optimization is proposed, an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation and obtains an improved success rate over strong baselines and ablations demonstrate that action chunks are especially impor...

Wen-Dong Li, J. Garcke · 0 citations
Preprint Aug 2026

NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments

This work presents NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking that encodes full-environment obstacle and boundary geometry in a target-relative vertex representation, allowing the policy to retain route-defining context throughout the approach.

Zihan Wang, Baixiang Huang, Yang Guan et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.