Safe Learning Under Irreversible Dynamics via Asking for Help
This work provides an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces and shows that this combination enables the agent to learn both safely and effectively.
Benjamin Plaut, Juan Liévano-Karim, Hanlin Zhu et al.
· 3 citations