Skip to content

Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories

Aug 2026 · Statistica Sinica · 0 citations
Mathematics Computer Science

TL;DR

A new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning is developed and a group-sparsity-based feature screening procedure is proposed that identifies, with high probability, a reduced feature set containing all relevant covariates.

Abstract

We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probability finite-sample error bounds for estimating the value function of a target policy and show that the bounds depend only logarithmically on the ambient dimension $d$, thereby alleviating the curse of dimensionality. In contrast to most existing theory for off-policy evaluation, which typically assumes access to many trajectories, our analysis guarantees accurate value estimation when either the number of trajectories or the time horizon is sufficiently large. In addition, we propose a group-sparsity-based feature screening procedure that identifies, with high probability, a reduced feature set containing all relevant covariates. Numerical experiments demonstrate the effectiveness of the proposed approach.

View source

Similar papers

#machine learning Preprint Sep 2026

Finite-Sample Theory for Fitted Q-Iteration When Actions Are Functions

Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire function, such as a fluence map in radiation therapy or a smooth movement trajectory in robotics. In this paper, we study the finite-sample theory for fitted Q-iteration (FQI) wi...

Ge-Fei Lin, Rui Miao, Xiao-Ke Zhang · 0 citations
#machine learning Preprint Sep 2026

Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations

A convex combination of one-step and two-step residuals with a fixed mixing weight is considered to provide a natural bias--variance trade-off between approximation error under a restricted value-function class and the increased variance arising from multi-step importance weighting.

Amitakshar Biswas, Yu-Han Li, Ruo-Qing Zhu · 0 citations
#machine learning Preprint Sep 2026

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

This paper proposes a model-based bootstrap framework for uncertainty quantification of OPE in finite-horizon, time-inhomogeneous Markov decision processes (MDPs), and establishes bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation for the target policy va...

Wei-Wei Wang, Yu-Qiang Li, Xian-Yi Wu et al. · 0 citations
Preprint Aug 2026

Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation

In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not ac...

Claire Chen, S. Liu, Licheng Luo et al. · 2 citations
#machine learning Preprint Sep 2026

Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, wh...

Suman Banerjee, Hiroyasu Tsukamoto · 0 citations
#machine learning Preprint Sep 2026

Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation

Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficienc...

Li-Can Kang, Jerry Zhijian Yang, Cheng Yuan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.