Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study
This work studies the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rather than by an expected discounted cost, and employs mini-batch transition risk mappings.
A. Patel, Andrzej Ruszczyński
· 0 citations