Skip to content

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

Sep 2026 · 0 citations
Computer Science

TL;DR

This work develops a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs), and validate the stabilising effect predicted by the theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.

Abstract

In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.

View source

Similar papers

Preprint Sep 2026

Is One Step Enough for Offline Policy Improvement?

This work studies how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy, and identifies improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and add...

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations
#machine learning Preprint Sep 2026

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

This paper proposes a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs), and establishes bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy va...

Wei-Wei Wang, Yu-Qiang Li, Xian-Yi Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead imp...

Hyukjun Yang, Jongchan Park, Narim Jeong et al. · 0 citations
#machine learning Preprint Sep 2026

Fast Regularized Policy Mirror Descent with One-Step TD Updates

Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we est...

Qi-Pei Chen, Wen-Ye Li, Yu-Le Sun et al. · 0 citations
#machine learning Preprint Oct 2026

Towards Optimal Policy Improvement

Practical Reinforcement Learning (RL) algorithms learn to solve Markov Decision Processes (MDPs) through iterative policy improvement in the presence of approximate evaluation. We study policy improvement from first principles, defining optimal policy improvement as producing the best policy attainable in a single upda...

Yaniv Oren, Viliam Vadocz, W. Zabka et al. · 0 citations
Preprint Sep 2026

Policy Gradient over History-Dependent Policy Classes for LQR with Domain Randomization

Domain Randomization (DR) has been widely used to overcome the sim-to-real gap by training a controller on a distribution of simulated environments via reinforcement learning. While DR can achieve robust performance simply using controllers synthesized via policy gradient (PG) methods, the optimization landscape is not...

Tesshu Fujinami, Bruce D. Lee, Anastasios Tsiamis et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.