Sep 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 12726-12733· 0 citations· 33 references
Computer Science
TL;DR
A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque–speed envelope and a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training.
Abstract
Achieving high-speed locomotion in quadrupedal robots remains highly challenging, as actuators operate near their physical limits and exhibit pronounced nonlinearities. However, many existing methods neglect actuator nonlinearities and physical constraints during training, leading to a significant sim-to-real gap under highly dynamic motions and limiting achievable performance. To address this issue, we propose a high-speed locomotion framework that reduces sim-to-real discrepancies and stabilizes learning over a wide command distribution. A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque–speed envelope. In addition, a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training. Experiments on the 36.5 kg quadruped BlackPanther2 (BP2) demonstrate speeds of up to $13.2\,\mathrm{m/s}$ on a treadmill and $11.65\,\mathrm{m/s}$ outdoors, establishing a new state-of-the-art and, to the best of our knowledge, a world record for quadrupedal robot locomotion. The results further highlight the importance of accurate actuator modeling in preventing non-physical policy exploitation, and show that ACS improves robustness without sacrificing performance. This paper was recommended for publication by Editor Clement Gosselin upon evaluation of the Associate Editor and Reviewers’ comments
Achieving biological-level running speeds has largely been pursued through advances in control algorithms, which improve the utilization of existing hardware. However, the ultimate speed limits remain governed by the underlying force and torque requirements of rapid locomotion, which are typically addressed through inc...
Yu-Cheng Tao, Yong-Bin Jin, Shao-wen Cheng et al.· 0 citations
Reinforcement learning (RL) has become a powerful tool for quadrupedal locomotion, and a sim-to-real approach is widely adopted to avoid hardware damage during training. However, the “sim-to-real gap” remains a critical challenge, particularly for robots driven by high-gear-ratio actuators, in which nonlinear friction...
Hansol Kang, Hyunyong Lee, Jiman Park et al.· Machines· 0 citations
A deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss that employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive ob...
Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo et al.· 1 citation
Reinforcement learning shows strong potential for bipedal locomotion control, while reduced-order models provide compact and interpretable physical priors. This paper presents an Angular Momentum Linear Inverted Pendulum (ALIP) guided reinforcement learning framework for dynamic locomotion of a small point-foot bipedal...
Yong-Ming Yue, Ying-Rong Chen, Yi-Xiao Zheng et al.· IEEE Robotics and Automation...· 0 citations
Experimental results demonstrate that the integrated system improves locomotion stability, energy efficiency, and terrain adaptability compared with baseline controllers, highlighting the effectiveness of combining a structured gait prior, lightweight residual coordination, and hardware-aware deployment for practical q...
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-...
Victor Paredes, Ayonga Hereid· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.