Skip to content
Preprint

NavPatch: Evidence-Guided Object-Level Costmap Correction with Vision-Language Models

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

NavPatch is presented, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model.

Abstract

Mobile robots typically rely on geometric maps for obstacle avoidance and path planning, but the resulting obstacle representation does not always match how an object should affect navigation. A low lying cable may be missed, a flexible curtain may create spurious blockage, and a traffic cone may require an exclusion region larger than its observed footprint. We present NavPatch, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model. Open vocabulary grounding localizes object instances, and LiDAR and RGB-D observations provide 3D support. Observation quality filtering and cross frame maintenance determine when each correction patch is committed, replaced, or revoked. In 50 real robot trials across five layouts, NavPatch achieves an overall success rate of 86.0%. An ablation study of four configurations with 200 runs in total shows that NavPatch improves the success rate from 70.0% to 86.0% and reduces the false commit rate from 68.4% to 40.7% compared with updates based only on the current observation.

View source

Similar papers

Preprint Aug 2026

OccPlanner: Goal-Aware Occupancy-Conditioned Diffusion Planner for PixelGoal Navigation

PixelGoal navigation specifies targets directly in the agent's camera view, providing a natural interface between high-level visual reasoning and low-level navigation. Depth can lift a visible target pixel into a metric PointGoal, but this estimate becomes unreliable under occlusion or sensor noise. Moreover, a PointGo...

Binling Huang, Nian-Jin Ye, Xi Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain whic...

Savio Cardoz, Santhiya Rajan · 0 citations
#artificial intelligence Preprint Sep 2026

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-re...

Guo-Qing Ma, Ming-Qi Yuan, Chen Gao et al. · 0 citations
Preprint Sep 2026

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interact...

Praveen Kumar, K. Guruprasad, Tushar Sandhan · 0 citations
Open access Aug 2026

DevGRU: Depth-Guided Visual Navigation Using a Collision-Aware Recurrent Model

The proposed DevGRU navigation system employs an action predictor that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles and has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines.

Kyung Min Han, Eunsom Kim, Young J. Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.