Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
Experiments on locomotion, navigation and manipulation show consistent gains in interaction efficiency, imitation performance, and robustness, indicating that RECON directs real-environment interaction toward recovery regions around the expert distribution that are underexplored by prior methods, and thereby learns a w...