Skip to content
Conference

Policy Gradient Optimization for Markov Decision Processes with Epistemic Uncertainty and General Loss Functions

· IISE Annual Conference & Expo 2025 · 0 citations

Abstract

Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty associated with unknown parameters, we take a Bayesian approach to estimate the parameters from data and impose a coherent risk functional (with respect to the Bayesian posterior distribution) on the general loss function. Since this formulation usually does not satisfy the interchangeability principle, it does not admit Bellman equations and cannot be solved by approaches based on dynamic programming. Therefore, we develop a policy gradient optimization approach to address this problem.  We utilize the dual representation of the coherent risk measure and extend the envelope theorem to derive the policy gradient. Our extension of the envelope theorem from the discrete case to the continuous case may be of independent interest. We then show the convergence of the proposed algorithm with a convergence rate of O(1/t), where t is the number of policy gradient iterations. We further extend our algorithm to an episodic setting, and establish the consistency of the extended algorithm and provide bounds on the number of iterations needed to achieve a constant error bound in each episode.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.