Skip to content
Preprint

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

Aug 2026 · 0 citations · 26 references
Mathematics

TL;DR

A novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control with the help of the Lyapunov-type operator's spectrum, which is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain.

Abstract

Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.

View source

Similar papers

Sep 2026

Carleman approximation based adaptive optimal control design of nonlinear systems: A three-phase policy iteration approach.

A novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration through Carleman linearization, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation.

Jian-Guo Zhao, Zhi-Jiang Gao, Chun-Yu Yang et al. · 0 citations
Open access Sep 2026

Reinforcement-Learning Robust Tracking Control of Discrete-Time Systems with Diagonal Scaling

This paper addresses a robust tracking problem for linear discrete-time systems by proposing a reinforcement learning (RL) control method based on a diagonal-scaling strategy, offering a solution tailored to the demands of enhanced reliability. To overcome a common limitation in policy-iteration-based RL design, namely...

Kan-Yang Jiang, Zheng Gao, Ye Zeng et al. · 0 citations
Sep 2026

Event-triggered reinforcement learning-based safe control for stochastic systems subject to asymmetric input constraints and unknown dynamics.

This paper investigates the safe optimal control (SOC) for input-constrained unknown stochastic systems via adaptive dynamic programming (ADP) and generalized fuzzy hyperbolic model (GFHM). Firstly, a GFHM is employed to approximate the unknown nonlinear terms of the stochastic system, thereby eliminating the need for...

Yu-Ling Liang, Feng-Lin Qin, Lei Liu et al. · 0 citations
Preprint Aug 2026

Iterative State- and Control-Dependent Model Predictive Control: A Jacobian-Free Formulation for Constrained Nonlinear Systems

This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex o...

Mohammadreza Kamaldar · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.