Skip to content
Preprint

AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention, which support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference.

Abstract

Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive policy stages instead of repeatedly injecting a terminal feature. Navigation-aware GFM memory retains historical VGGT global-attention KV states according to instruction relevance, geometric confidence, and transition novelty under a bounded per-layer budget. Retained states provide geometric context for subsequent observations before fusion with the policy. Across R2R-CE and RxR-CE, \method{} achieves strong performance using a single RGB stream without additional navigation-specific external data. Controlled ablations show that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations. Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention. These findings support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference. Code will be released upon acceptance at https://humanoid-research.github.io/adageovln/.

View source

Similar papers

Preprint Aug 2026

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

GaussVLA is proposed, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, and Depth-Aware Chain-of-Thought (DA-CoT) that performs structured, non-autoregressive geometric reasoning under language and flow-ti...

MD SELIM SAROWAR, Md Tanvir Islam, Sungho Kim et al. · 1 citation
Preprint Sep 2026

FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation

Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In particular, future obs...

Khang Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen et al. · 0 citations
Preprint Aug 2026

StrataVLA: Hierarchical and Efficient 3D Geometric Grounding for Vision-Language-Action Models

StrataVLA is introduced, a plug-and-play framework for hierarchical geometry grounding that achieves 98.53% average success on LIBERO suites while reducing geometry-model invocations by up to 88%, establishing hierarchical geometry injection as an effective and efficient way to achieve spatially grounded robotic contro...

Jin Cui, Zhao Pu, Bo Cai et al. · 0 citations
Preprint Sep 2026

NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory

Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces...

Kai Sheng, Liu-Yi Wang, Jin-Long Li et al. · 0 citations
Preprint Sep 2026

LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration

Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual...

Jian-He Zhao, Yan-Hua Qiu, Zhi-Yu Zhang et al. · 0 citations
Preprint Aug 2026

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

This work proposes TAMP-Nav, a unified framework for efficient embodied navigation that dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhanci...

Hongyan Feng, Sun-Lai Chen, Xuan-Yu Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.