Data-Driven Synthesis of Probabilistic Controlled Invariant Sets for Linear Markov Decision Processes
Abstract
We study data-driven computation of <italic>probabilistic controlled invariant sets</italic> (PCIS) for safety-critical reinforcement learning under unknown dynamics. Assuming a linear MDP model, we use regularized least squares and self-normalized confidence bounds to construct a conservative estimate of the states from which the system can be kept inside a prescribed safe region over an <inline-formula><tex-math notation="LaTeX">$N$</tex-math></inline-formula>-step horizon, together with the corresponding set-valued safe action map. This construction is obtained through a backward recursion and can be interpreted as a conservative approximation of the <inline-formula><tex-math notation="LaTeX">$N$</tex-math></inline-formula>-step safety predecessor operator. When the associated conservative-inclusion event holds, a conservative fixed point of the approximate recursion can be certified as an <inline-formula><tex-math notation="LaTeX">$(N,\epsilon)$</tex-math></inline-formula>-PCIS with confidence at least <inline-formula><tex-math notation="LaTeX">$\eta$</tex-math></inline-formula>. For continuous state spaces, we introduce a lattice abstraction and a Lipschitz-based discretization error bound to obtain a tractable approximation scheme. Finally, we use the resulting conservative fixed-point approximation as a runtime candidate PCIS in a practical shielding architecture with iterative updates, and illustrate the approach on a numerical experiment.