Skip to content

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

Sep 2026 · 0 citations
Computer Science

TL;DR

A dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance is proposed, and object-conditioned visual reasoning with conservative evidence qualification is introduced to improve observation reliability before spatial accumulation.

Abstract

Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.

View source

Similar papers

Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 0 citations
Preprint Sep 2026

Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation

SBFNav is introduced, a closed-loop navigation framework centered on a language- conditioned Spatial Belief Field (SBF) that preserves multiple spatial hypotheses under par- tial evidence and further confirms the advantages of spatial-belief modeling over single-point prediction.

Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al. · 0 citations
Preprint Sep 2026

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, li...

Jia-Jun Jiang, Chun-Liang Hua, Zi-Chun Chen et al. · 0 citations
Preprint Aug 2026

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor envi...

Deyi Zhu, Hao-Yu Fan, Yinan Zhu et al. · 2 citations
Preprint Aug 2026

Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation

DBFly is proposed, a vision-language waypoint prediction framework that introduces explicit vision-guided spatial deliberation before waypoint generation and develops a terminal-convergence-aware stopping strategy that characterizes terminal states through both target proximity and short-horizon motion convergence, ena...

Fanfu Xue, En Yu, Bo-Han Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

AGC-VLN (Air-Ground Collaborative VLN), the first training-free baseline for air-ground collaborative VLN, establishes that training-free methods decompose navigation into VLM-based semantic reasoning and deterministic geometric execution, exposing a collaboration interface.

Shu-Ning Zhang, Liang Li, Yun-Heng Wang et al. · 1 citation · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.