StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. In...