A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain.
Abstract
Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.
A novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration through Carleman linearization, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation.
Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al.· Neural Networks· 0 citations
This paper addresses a robust tracking problem for linear discrete-time systems by proposing a reinforcement learning (RL) control method based on a diagonal-scaling strategy, offering a solution tailored to the demands of enhanced reliability. To overcome a common limitation in policy-iteration-based RL design, namely...
Kan-Yang Jiang, Zheng Gao, Ye Zeng et al.· Eksploatacja I Niezawodnosc-...· 0 citations
The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence.
Karl Handwerker, Felix Thömmes, Lucas Günther et al.· 1 citation
This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems and proves that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations.
This paper investigates the safe optimal control (SOC) for input-constrained unknown stochastic systems via adaptive dynamic programming (ADP) and generalized fuzzy hyperbolic model (GFHM). Firstly, a GFHM is employed to approximate the unknown nonlinear terms of the stochastic system, thereby eliminating the need for...
Yu-Ling Liang, Feng-Lin Qin, Lei Liu et al.· Neural Networks· 0 citations
This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex o...
Mohammadreza Kamaldar· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.