Skip to content
Preprint

Data-Driven optimal control via Koopman operators and Hamilton-Jacobi-Bellman equations

Aug 2026 · 0 citations · 70 references
Mathematics

Abstract

This paper presents a data-driven stable manifold (DD-SM) method, which integrates Koopman operator representation learning with the geometric stable manifold approach to Hamilton-Jacobi-Bellman (HJB) equations, enabling end-to-end optimal feedback control synthesis from raw trajectory data without prior knowledge of system dynamics. We construct an augmented control system under a unified symmetric subspace decomposition (SSD) and extended dynamic mode decomposition (EDMD) framework for joint approximation of the drift field, control matrix and their spatial derivatives, and derive probabilistic finite-sample error bounds for invariant and non-invariant dictionary spaces to yield a provably accurate approximate characteristic system of HJB equation. Via Lyapunov-Perron operator and ODE perturbation analysis, we prove the data-driven stable manifold achieves monotonically decreasing semi-global error with growing training data. We further establish closed-loop exponential stability and quantify the optimality gap, both tightenable by refining model accuracy. An efficient algorithm pipeline with adaptive data generation and deep neural approximation is developed, outputting control signals within 1 millisecond. Experiments on a modified van der Pol oscillator verify the effectiveness of our method.

View source

Similar papers

Preprint Aug 2026

Iterative State- and Control-Dependent Model Predictive Control: A Jacobian-Free Formulation for Constrained Nonlinear Systems

This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex optimization with a sequence of constrained linear-quadratic programs. Refreezing the coefficient matrices along the previously predicted trajectory drives the iteration. Near the origin, we prove this sequence contracts to a unique fixed point. We explicitly bound the number of iterations required to reach any stopping tolerance, and we quantify the distance from the fixed point to a true Karush-Kuhn-Tucker point, showing this optimality gap vanishes quadratically as the state approaches the origin. Inflating the discrete algebraic Riccati equation generates terminal ingredients that guarantee recursive feasibility and asymptotic stability, even when the solver terminates early. We adapt the terminal penalty online, proving it remains uniformly bounded, and we secure output feedback through the block-observable canonical form, which extracts the exact system state directly from past inputs and outputs. Retaining the block-banded structure of the subproblem forces the computational cost to scale linearly with the horizon length $\ell$. This $O(\ell)$ complexity matches the iterative linear quadratic regulator (iLQR) but sharply undercuts the $O(\ell^3)$ scaling of dense sequential quadratic programming (SQP). Numerical studies on a saturated quadrotor, a nonholonomic integrator, and a nonminimum-phase plant illustrate the theoretical bounds and map how the algorithm compares with iLQR, SQP, and linear-parameter-varying MPC.

M. Kamaldar · 0 citations
Preprint Jul 2026

Bilinear Koopman-Based Robust Model Predictive Control for Unknown Nonlinear Systems via Contraction Metrics

A RMPC framework for unknown nonlinear systems with general nonlinear constraints based on data-driven bilinear Koopman realizations is proposed and robust satisfaction of the original nonlinear constraints is proved by the true closed-loop trajectory, recursive feasibility, and convergence to a neighborhood of the target state.

Yuki Higuchi, Kazuhiro Sato · 1 citation
Preprint Jul 2026

Trajectory-Regularized Stochastic Optimal Control via KL Divergence

We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theorem, the trajectory KL reduces to a quadratic drift mismatch penalty, yielding a modified running cost that preserves the dynamic programming (DP) structure. We derive the corresponding Hamilton--Jacobi--Bellman (HJB) equation and characterize the optimal policy. In the linear-quadratic (LQ) setting, the formulation admits a closed-form solution with an augmented control cost. Experiments show that the regularization parameter induces a trade-off between performance-driven and reference-preserving behavior, including cases with reference dynamics learned from offline data.

Mintae Kim, K. Sreenath · 0 citations
Preprint Jul 2026

Stochastic Stability of Nonlinear MPPI via Contraction Theory and Control Lyapunov Functions

Model Predictive Path Integral (MPPI) control is directly implementable on nonlinear systems because its online update requires only forward rollouts of the dynamics, not gradients, linearizations, or convex optimization. However, this algorithmic flexibility does not by itself provide a closed-loop stability certificate. This paper establishes such a certificate through a stability-inheritance argument. We assume that there exists a deterministic nonlinear MPC policy whose disturbance-free closed loop is certified by a Control Lyapunov Function terminal cost and a contraction metric, and we show that finite-sample MPPI inherits the nominal contraction when its sampling-based update approximates this reference policy with sufficient accuracy. The approximation error decomposes into a finite-temperature bias floor and a Monte Carlo term that vanishes at the inverse square-root rate in the sample count. Under an explicit small-gain condition, the resulting MPPI closed loop satisfies a finite-horizon, high-probability localized mean practical stability bound with residual floors due to MPPI approximation error, Gaussian process noise, and bad sampling events. The paper also gives an ISS-type restatement and a finite-horizon design procedure for choosing the localization set, temperature, and sample count.

Hyung-Jin Yoon, Hunmin Kim · 0 citations
Open access Aug 2026

Projection Neural Dynamics for Inverse Variational Inequality Problems: Stability Analysis and Applications to Sparse Signal Recovery

In this work, we develop a projection neural network based on a second-order dynamical model (SO-PDM) for solving inverse variational inequality problems (IVIPs) in Hilbert spaces. The proposed framework incorporates inertial and damping components, resulting in improved convergence behavior while ensuring feasibility through a projection operator. Under the Lipschitz continuity assumption on the operator, the proposed SO-PDM admits a unique global trajectory. Under the additional strong monotonicity assumption and suitable parameter conditions, convergence to the unique solution of the IVIP is established. A discrete-time formulation is derived via a finite-difference scheme, leading to a projection-based inertial algorithm with relaxation. Under suitable parameter conditions, the algorithm is shown to converge linearly to the unique solution of the IVIP, and under an additional parameter condition, the global asymptotic stability of the continuous-time SO-PDM is established via Lyapunov analysis. Furthermore, a numerical comparison in a higher-dimensional setting shows that the proposed algorithm converges faster and attains higher accuracy than the existing first-order projection method. Numerical experiments further confirm the effectiveness and stability of the proposed SO-PDM, including its application to sparse signal recovery in compressed sensing.

Vajahat Karim Khan, M. Sarfaraz, Hafiz Farooq Ahmad et al. · 0 citations