A dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance is proposed, and object-conditioned visual reasoning with conservative evidence qualification is introduced to improve observation reliability before spatial accumulation.
Abstract
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
SBFNav is introduced, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF) that preserves multiple spatial hypotheses under par- tial evidence and further confirms the advantages of spatial-belief modeling over single-point prediction.
Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al.· 0 citations
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, li...
Jia-Jun Jiang, Chun-Liang Hua, Zi-Chun Chen et al.· 0 citations
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor envi...
Deyi Zhu, Hao-Yu Fan, Yinan Zhu et al.· 2 citations
DBFly is proposed, a vision-language waypoint prediction framework that introduces explicit vision-guided spatial deliberation before waypoint generation and develops a terminal-convergence-aware stopping strategy that characterizes terminal states through both target proximity and short-horizon motion convergence, ena...
AGC-VLN (Air-Ground Collaborative VLN), the first training-free baseline for air-ground collaborative VLN, establishes that training-free methods decompose navigation into VLM-based semantic reasoning and deterministic geometric execution, exposing a collaboration interface.
Shu-Ning Zhang, Liang Li, Yun-Heng Wang et al.· 1 citation· ⚡1
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 30, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.