Skip to content
Open access

Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning in Long-Horizon AntMaze Navigation

Aug 2026 · Machines · 0 citations · 34 references

TL;DR

Results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved.

Abstract

Long-horizon navigation requires a high-level policy to select locally reachable subgoals, yet a scalar task reward provides little information about how an unsuitable proposal should be changed. We introduce Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning (SCS-HRL), a two-level method in which a topology- and clearance-aware programmatic supervisor evaluates each proposed subgoal and returns both a scalar score and a continuous target in the same subgoal space. The score trains the high-level critic, and the target enters a masked regression term for the high-level actor. Primitive actions are always conditioned on the actor’s subgoal; the supervisor is inactive during learned-policy evaluation. In AntMaze, using 6000 training episodes, five seeds, and 100 deterministic evaluation episodes per seed, SCS-HRL attained an 88.4±7.8% final success rate (mean ± sample standard deviation; 95% Student-t confidence interval [78.7%,98.1%]). The matched scalar-only condition and HIRO attained 0% rates. Applying the same route rule directly to the SCS-HRL low-level controllers yielded 82.2±9.9% success; the paired difference favored the learned high-level policy by 6.2 percentage points (95% confidence interval [2.1,10.3], p=0.013). Across three matched seeds, nonzero corrective weights of 0.5, 1.0, and 2.0 remained stable, whereas 0.25 was seed-sensitive. Term-level ablations further show that the continuous target, rather than the exact scalar-shaping formula, was the principal additional signal. Separate fixed-policy tests obtained 0% success rates on two unseen maze layouts. These results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic

An alternative reward shaping method (RS) is proposed that removes deceptive rewards at the expense of theoretical guarantees of PBRS, and another method named Locally-Guided Actor Critic (LG-AC) that rewards the agent for reaching intermediate goals achieves the best overall performance across tasks.

Olivier Serris, Stéphane Doncieux, Olivier Sigaud · 0 citations
Preprint Aug 2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al. · 7 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning

Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; how...

Heng-Rui Zhang, Yu-Hu Cheng, C. L. Philip Chen et al. · 1 citation
Preprint Sep 2026

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement...

Saksham Singh, Zhe-Yuan Hu, Max Sobol Mark et al. · 0 citations
#artificial intelligence Preprint Sep 2026

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier by executing each task's own verifier.

Jun-Yao Yang, Yu-Cheng Shi, Zhong-Zhi Li et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Reinforcement Learning with Decomposed Subtasks

Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feed...

Mattie Terzolo, Mikolaj Sacha, Ayan Sinha et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.