Skip to content
Open access

Assisted exploration of difficult transitions in humanoid motion tracking

Sep 2026 · Frontiers in Neurorobotics · 0 citations · 34 references

Abstract

Humanoid motion tracking provides a scalable interface for transferring kinematic human motion to whole-body robot control, but highly dynamic and contact-rich references remain difficult because underactuation and discontinuous contacts can restrict the closed-loop state distributions explored during training. External wrenches can reduce dynamic mismatch during imitation, but retaining such assistance in the final controller can compromise physical realism. This paper proposes residual-wrench-assisted reinforcement learning (RWARL), a two-stage framework that uses base residual wrenches as a temporary optimization scaffold. RWARL first solves a relaxed tracking problem with residual-wrench assistance and then anneals the residual-wrench strength while adapting early-termination bounds and segment sampling to observed training difficulty. To keep the policy adaptable as the training distribution shifts from assisted to unassisted dynamics, RWARL incorporates plasticity-preserving mechanisms during relaxation. Experiments on highly dynamic and contact-rich references show that RWARL generally improves global tracking accuracy, with the largest gains near difficult transitions where standard training tends to produce conservative behavior. Compared with the baseline, the final unassisted RWARL policies reduce global body-position tracking error by 8.6%–37.8% on the highly agile and contact-rich motions examined in this work and achieve the largest successful-frame-rate gain of 41.8 percentage points on jump twist . A closed-loop predecessor-state diagnostic further links these gains to more favorable predecessor-state distributions around critical phases. We deploy the final unassisted controllers on Unitree G1 hardware, demonstrating that the learned policies remain executable after residual assistance is removed. Together, these findings suggest that RWARL uses residual wrenches to reshape learning around difficult, low-reachability transitions, allowing the final unassisted policy to inherit better tracking behavior in phases that standard training tends to underexplore.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.