Skip to content
Preprint

Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

SBFNav is introduced, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF) that preserves multiple spatial hypotheses under par- tial evidence and further confirms the advantages of spatial-belief modeling over single-point prediction.

Abstract

Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or action, prematurely collapsing the spatial uncertainty inherent in incomplete evidence and ambiguous relations. To address this limitation, we introduce SBFNav, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF). Unlike ego-centric maps that primarily record what has been observed, SBF rep- resents a task-conditioned distribution over plausible target locations, preserving multiple spatial hypotheses under par- tial evidence. At each step, this distribution is updated from accumulated observations as new evidence becomes avail- able. Built on this representation, SBFNav selects the goal that best aligns with the instruction and observations as a met- ric waypoint for control. Experiments on both the original and revised CityNav benchmarks achieve the best reported overall performance. On the Test Unseen split, our method improves SR from 25.91% to 32.29% and SPL from 19.63% to 30.43%. Ablation studies further confirm the advantages of spatial-belief modeling over single-point prediction.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

A dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance is proposed, and object-conditioned visual reasoning with conservative evidence qualification is introduced to improve observation reliability before spatial accumulation.

Jian-Qiang Xiao, Xiang Deng, Yue-Xuan Sun et al. · 0 citations
Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 0 citations
Preprint Sep 2026

OmniNav: Robust Long-Horizon Target Navigation in Dynamic Environments

Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisions, and physical interactions. This requires three coupled capabilities: maintaining valid scene memory, revising target beliefs under partial observability, and selecting interaction-feasible navigation endpo...

Yu-Jie Tang, Mei-Ling Wang, Jin-Hao Jiang et al. · 0 citations
Oct 2026

ForexNav: Foresight Exploratory Navigation in Complex and Unknown Indoor Environments

Autonomous navigation in unknown, complex indoor environments remains challenging due to limited sensing range and severe partial observability. Conventional methods rely on local maps without foresight, causing dead-ends and long detours, while local goal selection based on Euclidean distance or frontier coverage fail...

Hong-Yu Song, Yun-Fang Ren, Ji-Gui Miao et al. · 0 citations
Preprint Aug 2026

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor envi...

Deyi Zhu, Hao-Yu Fan, Yinan Zhu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.