Sep 2026· International Conference on Photonic Computing, Algorithms, and Machine Vision· Vol 14320, pp. 143200Y - 143200Y-10· 0 citations· 14 references
Engineering
TL;DR
This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker with collaboration between a high-level manager and a low-level worker.
Abstract
Autonomous navigation of robots in complex indoor environments faces challenges such as long-term decision-making, sparse rewards, and low sample efficiency. Traditional end-to-end deep reinforcement learning methods often suffer from poor policy convergence and weak generalization ability when handling long-term tasks. This paper proposes a hierarchical reinforcement learning method based on an option framework. This method formalizes the navigation task as a semiMarkov decision process and constructs a two-layer decision architecture with collaboration between a high-level manager and a low-level worker. The core innovations are: proposing an option definition method based on semantic priors of the indoor environment; designing a hybrid reward mechanism that integrates sparse extrinsic rewards and dense intrinsic rewards; and adopting a three-stage course learning training process of “lower-level pre-training—higher-level training— joint fine-tuning”. Experimental results show that the proposed method achieves a task success rate of 67.0%, improves the average reward by approximately 45.2%, and reduces the number of interaction steps required for convergence by approximately 56.3%.
Developing effective robot navigation methods in crowded environments is essential for real-world applications. Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments, they often focus primarily on task-centric objectives and underrepresent social complian...
Takieddine Soualhi, Jacques Saraydaryan, Laëtitia Matignon· 0 citations
Mobile-robot navigation policies have typically assumed a fixed sensing input and robot platform. In this work, we investigate a teacher–student policy where teachers learn continuous velocity commands from LiDAR-based Twin Delayed Deep Deterministic Policy Gradient, and the student navigates using either LiDAR or came...
A. Amani, Sajjad Amani, AmirHossein Majidirad· Machines· 0 citations
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
Unmanned Aerial Vehicle (UAV) autonomous navigation is a key capability for UAVs, particularly when in complex and dynamic environments where continuous human control is impractical. While commonly used rule-based navigation and path-planning techniques can be effective in structured environments, they can be ineffecti...
B. S. Santhoshi, B. S, R. Shankar· 2026 International Conferenc...· 0 citations
A hierarchical reinforcement learning framework for autonomous highway driving that decomposes delayed-reward highway overtaking decision making into interpretable subtasks and achieves more reliable trap-escape performance than other hierarchical structures, including h-DQN and HIRO.
Zhihao Zhang, Ekim Yurtsever, K. Redmill· IEEE Access· 0 citations
Autonomous ecological monitoring requires robotic systems capable of efficiently exploring large natural environments while minimizing redundant traversal under limited communication and partial environmental knowledge. Although reinforcement learning provides adaptive decision making for multi-robot exploration, spars...
T. Gurunathan, A. Gangopadhyay· Frontiers in Robotics and AI· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026