Skip to content

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

Aug 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP · 0 citations
Medicine

TL;DR

A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.

Abstract

Vision-and-Language Navigation in Continuous Environments (VLN-CE) has emerged as a pivotal challenge in Embodied AI, requiring an agent to navigate 3D spaces guided by natural language instructions. Drawing inspiration from human cognition, world models provide a powerful paradigm by predicting environment dynamics and enabling reasoning beyond immediate observations. However, existing world model-based VLN methods remain static once trained - their representations rely on fixed correlationbased priors rather than adaptive causal structures, making them unable to accommodate evolving confounders and changing observation-action dependencies across environments. This rigidity leads to overfitting to training-specific patterns and degraded performance under distribution shifts. To address this limitation, we propose a causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes. Our model learns unified latent states that integrate vision, language, and action, while addressing spurious correlations through a dual-level intervention mechanism: at the observation level, frequency-domain perturbations simulate superficial appearance variations to enhance perceptual robustness; at the representation level, cross-episode confounder buffers perform counterfactual substitution to approximate the influence of latent confounding factors. Beyond static world modeling, our framework continuously evolves, refining these proxy representations across episodes, enabling efficient adaptation to previously unseen environments. Building on this evolving causally-inspired foundation, our world model supports counterfactual reasoning and strengthens generalization across diverse navigation contexts. Extensive evaluations on established VLN-CE benchmarks demonstrate that our method outperforms existing approaches, delivering superior navigation performance across diverse scenarios. Real-world robot evaluations further validate the practicality of our approach. Code is available in the Supplementary Material.

View source

Similar papers

Preprint Aug 2026

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

AWM-VLA is presented, a unified framework that embeds aligned world modeling directly inside a diffusion-transformer policy and is compatible with any diffusion or flow-matching policy, making aligned world modeling an inexpensive, broadly applicable component of generalist manipulation.

Lan-Ji An, Da-Wei Liu, Jin Li et al. · 0 citations
Review Sep 2026

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enab...

Shaunak A. Mehta, Ananya Hazarika, Haochen Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

This work proposes BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations, and introduces an asynchronous rectified-flow inference strategy wit...

Bing Zhan, Shu-Yao Shang, Shuo Lu et al. · 1 citation
Preprint Sep 2026

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.

Yi-Cheng Jiang, Ze-Sen Gan, Xiao-Bo Wang et al. · 0 citations
Preprint Sep 2026

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

This work uses intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions and shows that these navigation policies are sensitive to all input modalities and do not depend on a single one.

Débora Oliveira Makowski, Samiran Gode, Abhijeet Nayak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.