SBFNav is introduced, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF) that preserves multiple spatial hypotheses under par- tial evidence and further confirms the advantages of spatial-belief modeling over single-point prediction.
Abstract
Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or action, prematurely collapsing the spatial uncertainty inherent in incomplete evidence and ambiguous relations. To address this limitation, we introduce SBFNav, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF). Unlike ego-centric maps that primarily record what has been observed, SBF rep- resents a task-conditioned distribution over plausible target locations, preserving multiple spatial hypotheses under par- tial evidence. At each step, this distribution is updated from accumulated observations as new evidence becomes avail- able. Built on this representation, SBFNav selects the goal that best aligns with the instruction and observations as a met- ric waypoint for control. Experiments on both the original and revised CityNav benchmarks achieve the best reported overall performance. On the Test Unseen split, our method improves SR from 25.91% to 32.29% and SPL from 19.63% to 30.43%. Ablation studies further confirm the advantages of spatial-belief modeling over single-point prediction.
A dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance is proposed, and object-conditioned visual reasoning with conservative evidence qualification is introduced to improve observation reliability before spatial accumulation.
Jian-Qiang Xiao, Xiang Deng, Yue-Xuan Sun et al.· 0 citations
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisions, and physical interactions. This requires three coupled capabilities: maintaining valid scene memory, revising target beliefs under partial observability, and selecting interaction-feasible navigation endpo...
Yu-Jie Tang, Mei-Ling Wang, Jin-Hao Jiang et al.· 0 citations
Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fail...
Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al.· IEEE Robotics and Automation...· 0 citations
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor envi...
Deyi Zhu, Hao-Yu Fan, Yinan Zhu et al.· 2 citations
This work introduces RiverVLN, to its knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs.
Jie-Ling Wu, Yue-Hao Huang, Jia-Jun Lv et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.