Skip to content
Preprint

Policy Iteration for Domain Randomized Linear Quadratic Systems

Sep 2026 · 0 citations · 28 references
Mathematics Computer Science Engineering

Abstract

In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability across all sampled systems at each iteration. We show that the method yields monotonic improvement of the sample-average objective and that a stabilizing step size always exists. Under standard smoothness assumptions, the iterates converge subsequentially to stationary points, and under a gradient-dominance condition, we obtain a global linear convergence rate.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.