Finite-sample bounds are derived that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality.
Yu-Ling Jiao, Li-Can Kang, Jerry Zhijian Yang et al.· 0 citations
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample utilization efficienc...
Li-Can Kang, Jerry Zhijian Yang, Cheng Yuan et al.· 0 citations
In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. T...
Xiao-Hong Chen, Yu-Ling Jiao, Li-Can Kang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.