Oct 2025· IEEE Transactions on robotics· Vol 42, pp. 3469-3488· 10 citations· 62 references
Computer Science
TL;DR
This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level, achieving higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen gaits and uneven terrain.
Abstract
Model predictive control (MPC) provides interpretable, tunable locomotion controllers grounded in physical models, but its robustness depends on frequent replanning and is limited by model mismatch and real-time computational constraints. Reinforcement learning (RL), by contrast, can produce highly robust behaviors through stochastic training but often lacks interpretability, suffers from out-of-distribution failures, and requires intensive reward engineering. This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level. We develop a kinodynamic whole-body MPC formulation evaluated across thousands of agents in parallel at 100 Hz for RL training. The residual policy learns to make targeted corrections to the MPC outputs, combining the interpretability and constraint handling of model-based control with the adaptability of RL. The model-based control prior acts as a strong bias, initializing and guiding the policy toward desirable behavior with a simple set of rewards. Compared to standalone MPC or end-to-end RL, our approach achieves higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen gaits and uneven terrain.
A framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions is proposed.
Emek Barış Küçüktabak, Karankumar Patel, Zhao-Dong Yang et al.· 3 citations
This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.
Martin Schuck, Maks Sorokin, S. Manni et al.· 1 citation
A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.
This work introduces Gradient Control Rewards (GCR), an interpretable, control-inspired reward-design methodology for accelerating agent training by modulating the reward signal based on the temporal dynamics of system error, Inspired by classical control theory.
Anas Aburaya, H. Selamat, M. Muslim et al.· IEEE Access· 0 citations
Solver-Gradient Guided Reinforcement Learning is proposed, a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation that reaches PPO's best closed-loop return with up to 70.6% fewer samples, and outperforms GB-PL baselines by at least 54% in closed-loop return.
Baha Zarrouki, Arslan Thobani, Jasper Hoffmann et al.· 0 citations
Shape-Aware Reinforcement Learned Model Predictive Control is proposed, a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification that preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL.
Rui-Hua Han, Rui Gao, Zhe Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.