Preprint
Jul 2026
Joint On-and-Off Policy Learning for Vision-and-Language Navigation
JOP-VLN is introduced, a novel VLN framework that synergistically combines off-policy imitation learning and on-policy exploration within a three-stage training pipeline, featuring high-entropy trajectory sampling to enhance RL training efficiency and an error-correction-prioritized trajectory sorting strategy for effective error correction.
Qin He, Lingqing Zhao, Kevin Zheng et al.
· 0 citations