Model-free Reinforcement Learning for Continuous Time and State: A Stochastic Maximum Principle Approach
This paper develops a model-free reinforcement learning (RL) algorithm based on the stochastic maximum principle for continuous-time stochastic control problems with continuous state and action spaces. For a parameterized Markovian policy, we establish the existence of the decoupling field for the adjoint backward stoc...