Skip to content

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

Sep 2026 · 0 citations · 36 references
Mathematics Computer Science

TL;DR

This paper proposes a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs), and establishes bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value.

Abstract

Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essential for safe and risk-aware decision-making. A comprehensive way to unify these tasks is to estimate the sampling distribution of the evaluation error. Existing approaches, however, often suffer from limited robustness, scalability, or finite-sample validity. In this paper, we propose a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs). Unlike classical bootstrap methods that rely on resampling complete episodes, the proposed method regenerates trajectories from an estimated MDP and can therefore accommodate a much broader range of offline data formats, including complete trajectories, transition-level observations, and trajectory fragments. This flexibility further improves finite-sample statistical efficiency. We establish bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy value. Extensive simulations show that the proposed method accurately captures the sampling distribution of the OPE estimator, yielding tighter confidence intervals and more accurate variance estimates in most settings.

View source

Similar papers

Preprint Aug 2026

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not ac...

Claire Chen, S. Liu, Licheng Luo et al. · 2 citations
#machine learning Preprint Sep 2026

From Static Policies to Adaptive Priors in Offline Reinforcement Learning

This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction.

Tian-Wei Ni, Vineet Jain, Akash Karthikeyan et al. · 1 citation
#machine learning Preprint Sep 2026

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

This work develops a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs), and validate the stabilising effect predicted by the theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpar...

Nam Phuong Tran, Trinh Ha Mai Huynh, T. P. Le et al. · 0 citations
Preprint Sep 2026

Is One Step Enough for Offline Policy Improvement?

This work studies how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy, and identifies improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and add...

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations
#artificial intelligence Conference Sep 2026

Online Robust Reinforcement Learning Through Monte-Carlo Planning

A new robust variant of MCTS that mitigates dynamical model ambiguities to bridge the gap between simulation-based planning and real-world deployment and empirical evidence is provided that this method achieves robust performance in planning problems even under significant ambiguity in the underlying reward distributio...

T. Dam, Kishan Panaganti, Brahim Driss et al. · 4 citations

Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories

A new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning is developed and a group-sparsity-based feature screening procedure is proposed that identifies, with high probability, a reduced feature set containing all relevant covariates.

Tuo-Yi Zhao, Cheng-Chun Shi, Zheng-Ling Qi et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.