This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems and proves that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations.
Abstract
Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and backward costate propagation. This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems. We derive the policy-gradient expression where the state and costate rollouts required for policy-gradient evaluation are formulated as residual-minimization problems and solved using Gauss-Newton (GN) iterations with parallel associative scans. For closed-loop systems that are globally asymptotically stable and locally exponentially stable, we show that the residual-minimization problems satisfy a local Polyak-Lojasiewicz (PL) inequality and that the GN iterates converge locally at a quadratic rate. Moreover, the PL constant, the size of the convergence neighborhood, and the quadratic convergence bound are independent of the rollout horizon T. We also prove that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations. Finally, an inertia-wheel pendulum example with interconnection and damping assignment passivity-based control (IDA-PBC) demonstrates improved closed-loop performance and the computational benefits of the proposed parallel policy-gradient framework.
A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtai...
The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence.
Karl Handwerker, Felix Thömmes, Lucas Günther et al.· 1 citation
A novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration through Carleman linearization, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation.
Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al.· Neural Networks· 0 citations
This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex o...
This paper develops a model-free reinforcement learning (RL) algorithm based on the stochastic maximum principle for continuous-time stochastic control problems with continuous state and action spaces. For a parameterized Markovian policy, we establish the existence of the decoupling field for the adjoint backward stoc...
In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability acros...
Abbas Pasdar, F. Yaghmaie· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.