Skip to content
Open access

Residual MPC: Blending Reinforcement Learning With GPU-Parallelized Model Predictive Control

Oct 2025 · IEEE Transactions on robotics · Vol 42, pp. 3469-3488 · 10 citations · 62 references
Computer Science

TL;DR

This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level, achieving higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen gaits and uneven terrain.

Abstract

Model predictive control (MPC) provides interpretable, tunable locomotion controllers grounded in physical models, but its robustness depends on frequent replanning and is limited by model mismatch and real-time computational constraints. Reinforcement learning (RL), by contrast, can produce highly robust behaviors through stochastic training but often lacks interpretability, suffers from out-of-distribution failures, and requires intensive reward engineering. This work presents a GPU-parallelized residual architecture that tightly integrates MPC and RL by blending their outputs at the torque-control level. We develop a kinodynamic whole-body MPC formulation evaluated across thousands of agents in parallel at 100 Hz for RL training. The residual policy learns to make targeted corrections to the MPC outputs, combining the interpretability and constraint handling of model-based control with the adaptability of RL. The model-based control prior acts as a strong bias, initializing and guiding the policy toward desirable behavior with a simple set of rewards. Compared to standalone MPC or end-to-end RL, our approach achieves higher sample efficiency, converges to greater asymptotic rewards, expands the range of trackable velocity commands, and enables zero-shot adaptation to unseen gaits and uneven terrain.

Read PDF

Similar papers

Preprint Sep 2026

Real-World Reinforcement Learning with MPC Scaffolding for Dexterous Manipulation

A framework that uses sampling-based model predictive control (MPC) as scaffolding for real-world dexterous RL, providing structured prior experience and task-directed guidance during learning without human demonstrations or corrective actions is proposed.

Emek Barış Küçüktabak, Karankumar Patel, Zhao-Dong Yang et al. · 3 citations
Preprint Aug 2026

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.

Martin Schuck, Maks Sorokin, S. Manni et al. · 1 citation
#reinforcement learning Preprint Aug 2026

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

A novel framework is proposed that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes and achieves higher returns and faster convergence.

Hossein Abdi, S. Dash, Ming-Fei Sun · 0 citations
Open access 2026

GCR-RL: Gradient Control Reward Shaping for Reinforcement Learning

This work introduces Gradient Control Rewards (GCR), an interpretable, control-inspired reward-design methodology for accelerating agent training by modulating the reward signal based on the temporal dynamics of system error, Inspired by classical control theory.

Anas Aburaya, H. Selamat, M. Muslim et al. · 0 citations
#machine learning Preprint Sep 2026

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

Solver-Gradient Guided Reinforcement Learning is proposed, a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation that reaches PPO's best closed-loop return with up to 70.6% fewer samples, and outperforms GB-PL baselines by at least 54% in closed-loop return.

Baha Zarrouki, Arslan Thobani, Jasper Hoffmann et al. · 0 citations
Preprint Aug 2026

SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

Shape-Aware Reinforcement Learned Model Predictive Control is proposed, a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification that preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL.

Rui-Hua Han, Rui Gao, Zhe Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.