Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unr...
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv· 0 citations
Legged robots have demonstrated a remarkable ability to traverse various terrains, yet generating effective loco-manipulation behaviors remains challenging. A key difficulty is that object and terrain parameters are typically unknown to the robot, and mismatches between these parameters and their simulated counterparts...
Res-HIL is introduced, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy that improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations.
M. Iavorskaia, C. Dietz, Sebastian Albrecht et al.· 0 citations
This work freezes the visual encoder and confine set propagation to a low-dimensional interface between it and the downstream policy, with the interface set calibrated from held-out camera-pose perturbations to achieve reachability analysis for visuomotor policies.
Yan-Liang Huang, Zhuo-Cheng Zhang, Peng Xie et al.· 0 citations
The resulting algorithm reduces training-time falls by factors of 233x, 48x, and 26x on HalfCheetah, Ant, and Unitree Go1 over standard PPO, while matching or exceeding PPO's final reward, and on Ant, where the recovery policy is unreliable, it is the only method that reaches 80% of the best final reward.
E. Daneshmand, Majid Khadiv, Glen Berseth et al.· arXiv.org· 0 citations
This paper proposes a nested kino-dynamic framework for rapid feasibility checking and dynamically consistent trajectory generation given a candidate contact sequence and shows that the generated trajectories can be tracked using a reinforcement learning (RL)-based controller and are of sufficiently high quality for ex...
Michal Ciebielski, Shafeef Omar, Aaron M. Johnson et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.