The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence.
Abstract
This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fr\'echet derivative of the sequential PI map at a Nash equilibrium. The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence. Since finding an initial stabilizing solution is a major challenge in policy iteration, we also propose a homotopy-based initialization that ensures a valid starting point. The effectiveness of the proposed PI algorithm and the analytical results are verified through a numerical example.
This paper presents a method to solve the inverse problem for N-player infinite-horizon linear-quadratic (LQ) differential games with state- and control-dependent noise. For this stochastic setting, we derive necessary and sufficient conditions for linear feedback Nash equilibria, which take the form of coupled stochas...
Lucas Günther, Karl Handwerker, Felix Thömmes et al.· 1 citation
A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtai...
This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems and proves that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations.
In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability acros...
A novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration through Carleman linearization, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation.
Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al.· Neural Networks· 0 citations
This paper develops a critic-free policy iteration (PI) method for continuous-time linear zero-sum games. The central idea is to characterize the saddle-point policies directly in the joint policy space, rather than treating the quadratic value matrix as an iterative variable. A policy game Riccati equation (PGRE) is i...
Jia-Cheng Wu, Yang Zhu, Hong-Ye Su· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.