Skip to content

A hierarchical reinforcement learning approach for robot navigation integrating LiDAR priors and the options framework

Sep 2026 · International Conference on Photonic Computing, Algorithms, and Machine Vision · Vol 14320, pp. 143200Y - 143200Y-10 · 0 citations · 14 references
Engineering

TL;DR

This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker with collaboration between a high-level manager and a low-level worker.

Abstract

Autonomous navigation of robots in complex indoor environments faces challenges such as long-term decision-making, sparse rewards, and low sample efficiency. Traditional end-to-end deep reinforcement learning methods often suffer from poor policy convergence and weak generalization ability when handling long-term tasks. This paper proposes a hierarchical reinforcement learning method based on an option framework. This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker. The core innovations are: proposing an option definition method based on semantic priors of the indoor environment; designing a hybrid reward mechanism that integrates sparse extrinsic rewards and dense intrinsic rewards; and adopting a three-stage course learning training process of “lower-level pre-training—higher-level training— joint fine-tuning”. Experimental results show that the proposed method achieves a task success rate of 67.0%, improves the average reward by approximately 45.2%, and reduces the number of interaction steps required for convergence by approximately 56.3%.

View source

Similar papers

Preprint Aug 2026

Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling

Developing effective robot navigation methods in crowded environments is essential for real-world applications. Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments, they often focus primarily on task-centric objectives and underrepresent social complian...

Takieddine Soualhi, Jacques Saraydaryan, Laëtitia Matignon · 0 citations
Open access Aug 2026

A Unified Conditional Policy for Multi-Robot Navigation via LiDAR-to-Vision Distillation

Mobile-robot navigation policies have typically assumed a fixed sensing input and robot platform. In this work, we investigate a teacher–student policy where teachers learn continuous velocity commands from LiDAR-based Twin Delayed Deep Deterministic Policy Gradient, and the student navigates using either LiDAR or came...

A. Amani, Sajjad Amani, AmirHossein Majidirad · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations
Conference Aug 2026

A Multimodal Deep Reinforcement Learning Framework for Autonomous UAV Navigation in Gazebo–ROS Environments

Unmanned Aerial Vehicle (UAV) autonomous navigation is a key capability for UAVs, particularly when in complex and dynamic environments where continuous human control is impractical. While commonly used rule-based navigation and path-planning techniques can be effective in structured environments, they can be ineffecti...

B. S. Santhoshi, B. S, R. Shankar · 0 citations
Open access 2026

Extensive Exploration in Highway Overtaking Scenarios Using Hierarchical Reinforcement Learning

A hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks and achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.

Zhihao Zhang, Ekim Yurtsever, K. Redmill · 0 citations
Review Open access Sep 2026

Geometric priors for reinforcement learning-based multi-robot ecological monitoring

Autonomous ecological monitoring requires robotic systems capable of efficiently exploring large natural environments while minimizing redundant traversal under limited communication and partial environmental knowledge. Although reinforcement learning provides adaptive decision making for multi-robot exploration, spars...

T. Gurunathan, A. Gangopadhyay · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.