Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. In...
Anh Dao, Q. Phạm, Le Danh Vinh et al.· 0 citations
Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention, which support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference.
Q. Phạm, Anh Dao, Danh Vinh Le et al.· 0 citations
Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task-relevant hypotheses from the task-time interface. We present OpenBelief-Nav, an ev...
Dinh Tuan Nguyen, Anh Dao, Phuong Nam Dang et al.· 0 citations
HumanoidVLN, a physics-grounded simulator and benchmark for VLN across diverse humanoid embodiments, and compatibility with NaVILA, DualVLN, StreamVLN, and JanusVLN are presented.
Q. Phạm, Anh Dao, T. Nguyen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.