Jul 2026· Fall Joint Computer Conference· pp. 73-80· 0 citations· 16 references
Abstract
Cloud-edge collaborative computing enables resource-constrained robots to leverage powerful cloudhosted models while retaining real-time on-device perception and control. Vision-Language Navigation in Continuous Environments (VLN-CE), which requires a robot to follow natural-language instructions through complex scenes, is a representative task that benefits from this paradigm: it relies on large Visual Language Models (VLMs) for multimodal reasoning yet demands responsive execution at the edge. However, existing VLM-based approaches remain constrained by limited context windows and insufficient planning capabilities for long-horizon tasks. We present EntityNav, an entity-centric stepwise planning framework for VLN-CE designed for cloudedge deployment. EntityNav comprises two integrated modules executed on the cloud: (1) Entity-Guided Stepwise Language Planning, which decomposes instructions into sequential, entity-centered sub-goals for explicit progress tracking, and (2) Entity-Aware Chain-of-Thought Reasoning, which generates a multi-stage structured reasoning chain whose hidden-state representations directly condition the action prediction head, regularized by a reasoning-action consistency loss. On the robot side, an edge-level module performs real-time visual capture and local trajectory refinement, with asynchronous communication overlapping cloud inference and physical motion to preserve responsiveness; an edge-side fallback mechanism further maintains safe navigation during transient cloud delays. Experiments on R2R-CE and RxR-CE benchmarks show that EntityNav achieves success rates of 62.7% and 60.3% respectively, demonstrating competitive performance against baselines. Real-world deployment on a quadruped robot further shows the framework's effectiveness under practical cloud-edge conditions.
Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowled...
This survey revisits VLN from a component-internal perspective, viewing it as a navigation system composed of internal components such as instructions, environment representations, and embodied agents, and organizing existing work around the functions and interactions of these components.
Capek 0.5 is presented, an embodied vision-language model built around an execution-centric capability taxonomy that improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied ta...
Ying Chen, Wei-Zhen Li, Zhe Hu et al.· 1 citation· ⚡1
This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.
Howard Lu, Shalfun Li, Porter Pan et al.· 0 citations
Large language models (LLMs) are increasingly used as natural-language interfaces for robotic systems, yet their integration with Robot Operating System (ROS)-based navigation remains limited by two gaps. First, navigation data such as occupancy grids are represented as raw geometric messages that are difficult for LLM...
Jungsoo Lee, Jaegyun Park, Wansoo Kim· 0 citations
This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...
Ze-Yuan Ma, Jiaxin Chen, Di Huang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.