Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning
This work proposes a fundamental reformulation of the RL objective by replacing the scalar reward with a distribution over reward functions, and applying a non-linear objective over sets of actions, and develops a framework in which calibrated behavioural diversity emerges naturally, remains controllable through the re...