Open access
Aug 2026
Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning in Long-Horizon AntMaze Navigation
Results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved.
Li-Dong Sun, Ye Wang, Zhennan Fan et al.
· Machines · 0 citations