Open access
Jul 2021
Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
A novel HRL model is proposed that supports direct off-policy correction based on a Flow-based Deep Generative Model (FDGM) that leverages the inverse operation of FDGM to achieve goals aligned with the current knowledge of the lower-level policy.
Jaeyoon Kim, Junyu Xuan, C. Liang et al.
· Journal of Artificial Intell... · 0 citations