Skip to content
Preprint

OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

OptiSight is proposed, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture that demonstrates reliable zero-shot navigation across diverse indoor scenarios, including obstacle avoidance and semantic ambiguity, while operating within an 8~GB VRAM budget.

Abstract

Autonomous indoor navigation requires both semantic understanding and precise geometric control. We propose OptiSight, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture. Grounded-SAM localizes open-vocabulary targets, while camera projection geometry converts visual observations into navigation commands without requiring dense mapping. The VLM is queried only at key decision points, reducing computational overhead while geometric control handles continuous navigation. Experiments in AI Habitat demonstrate reliable zero-shot navigation across diverse indoor scenarios, including obstacle avoidance and semantic ambiguity, while operating within an 8~GB VRAM budget. The source code is available at https://github.com/avanalperen/OptiSight-Python-Multimodal-CoT-for-Visual-Reasoning.

View source

Similar papers

Conference Aug 2026

AVLF-Nav: Embodied Object Navigation via Active Visual Exploration and Lightweight Vision–Language Fusion

Embodied object navigation must couple partial visual perception with efficient exploration in unseen indoor environments. Existing agents often align language and vision without explicitly deciding which viewpoint is most informative, causing repeated visits and inefficient trajectories. This paper proposes AVLF-Nav,...

Nan Wang · 0 citations
Preprint Sep 2026

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...

Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 3 citations
Preprint Sep 2026

NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route

NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy, is presented, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.

Ayun Lee, Jiseon Kim, Giseop Kim · 0 citations
Preprint Sep 2026

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interact...

Praveen Kumar, K. Guruprasad, Tushar Sandhan · 0 citations
Preprint Sep 2026

VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

A VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches that preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference.

Enrico Saccon, Tommaso Faraci, Iñigo De La Ossa Zarzuelo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.