Skip to content
Preprint

Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers

Sep 2026 · 0 citations · 17 references
Engineering Computer Science Mathematics

TL;DR

This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems and proves that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations.

Abstract

Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and backward costate propagation. This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems. We derive the policy-gradient expression where the state and costate rollouts required for policy-gradient evaluation are formulated as residual-minimization problems and solved using Gauss-Newton (GN) iterations with parallel associative scans. For closed-loop systems that are globally asymptotically stable and locally exponentially stable, we show that the residual-minimization problems satisfy a local Polyak-Lojasiewicz (PL) inequality and that the GN iterates converge locally at a quadratic rate. Moreover, the PL constant, the size of the convergence neighborhood, and the quadratic convergence bound are independent of the rollout horizon T. We also prove that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations. Finally, an inertia-wheel pendulum example with interconnection and damping assignment passivity-based control (IDA-PBC) demonstrates improved closed-loop performance and the computational benefits of the proposed parallel policy-gradient framework.

View source

Similar papers

Preprint Aug 2026

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtai...

Xinyu Cao, Bing-Chang Wang, Ying Cao · 0 citations
Sep 2026

Carleman approximation based adaptive optimal control design of nonlinear systems: A three-phase policy iteration approach.

A novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration through Carleman linearization, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation.

Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al. · 0 citations
Preprint Aug 2026

Iterative State- and Control-Dependent Model Predictive Control: A Jacobian-Free Formulation for Constrained Nonlinear Systems

This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex o...

Mohammadreza Kamaldar · 1 citation
Preprint Sep 2026

Model-free Reinforcement Learning for Continuous Time and State: A Stochastic Maximum Principle Approach

This paper develops a model-free reinforcement learning (RL) algorithm based on the stochastic maximum principle for continuous-time stochastic control problems with continuous state and action spaces. For a parameterized Markovian policy, we establish the existence of the decoupling field for the adjoint backward stoc...

Li-Jun Bo, Yi-Jie Huang, Jing-Fei Wang · 0 citations
Preprint Sep 2026

Policy Iteration for Domain Randomized Linear Quadratic Systems

In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability acros...

Abbas Pasdar, F. Yaghmaie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.