Skip to content
Conference

EntityNav: an Entity-Centric Stepwise Planning Framework for Vision-Language Navigation

Jul 2026 · Fall Joint Computer Conference · pp. 73-80 · 0 citations · 16 references

Abstract

Cloud-edge collaborative computing enables resource-constrained robots to leverage powerful cloudhosted models while retaining real-time on-device perception and control. Vision-Language Navigation in Continuous Environments (VLN-CE), which requires a robot to follow natural-language instructions through complex scenes, is a representative task that benefits from this paradigm: it relies on large Visual Language Models (VLMs) for multimodal reasoning yet demands responsive execution at the edge. However, existing VLM-based approaches remain constrained by limited context windows and insufficient planning capabilities for long-horizon tasks. We present EntityNav, an entity-centric stepwise planning framework for VLN-CE designed for cloudedge deployment. EntityNav comprises two integrated modules executed on the cloud: (1) Entity-Guided Stepwise Language Planning, which decomposes instructions into sequential, entity-centered sub-goals for explicit progress tracking, and (2) Entity-Aware Chain-of-Thought Reasoning, which generates a multi-stage structured reasoning chain whose hidden-state representations directly condition the action prediction head, regularized by a reasoning-action consistency loss. On the robot side, an edge-level module performs real-time visual capture and local trajectory refinement, with asynchronous communication overlapping cloud inference and physical motion to preserve responsiveness; an edge-side fallback mechanism further maintains safe navigation during transient cloud delays. Experiments on R2R-CE and RxR-CE benchmarks show that EntityNav achieves success rates of 62.7% and 60.3% respectively, demonstrating competitive performance against baselines. Real-world deployment on a quadruped robot further shows the framework's effectiveness under practical cloud-edge conditions.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowled...

Yun-Zhe Xu, Zhe Liu · 0 citations
Review Open access Aug 2026

Vision-and-Language Navigation: A Component-Centric Survey of Interactions, Coupling, and Deployment

This survey revisits VLN from a component-internal perspective, viewing it as a navigation system composed of internal components such as instructions, environment representations, and embodied agents, and organizing existing work around the functions and interactions of these components.

Xiang-Xun Wu, Yin-Sheng Wu, Xiaojiang Peng · 0 citations
Preprint Aug 2026

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Capek 0.5 is presented, an embodied vision-language model built around an execution-centric capability taxonomy that improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied ta...

Ying Chen, Wei-Zhen Li, Zhe Hu et al. · 1 citation · ⚡1
Preprint Sep 2026

X-Planner: Event-Structured Task Planning for Embodied Intelligence

This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.

Howard Lu, Shalfun Li, Porter Pan et al. · 0 citations
Preprint Sep 2026

Spatial and Semantic Reasoning for LLM-Driven Robot Navigation via MCP

Large language models (LLMs) are increasingly used as natural-language interfaces for robotic systems, yet their integration with Robot Operating System (ROS)-based navigation remains limited by two gaps. First, navigation data such as occupancy grids are represented as raw geometric messages that are difficult for LLM...

Jungsoo Lee, Jaegyun Park, Wansoo Kim · 0 citations
Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.