Aug 2026· IEEE Transactions on Pattern Analysis and Machine Intelligence· Vol PP· 0 citations
Medicine
TL;DR
A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.
Abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) has emerged as a pivotal challenge in Embodied AI, requiring an agent to navigate 3D spaces guided by natural language instructions. Drawing inspiration from human cognition, world models provide a powerful paradigm by predicting environment dynamics and enabling reasoning beyond immediate observations. However, existing world model-based VLN methods remain static once trained - their representations rely on fixed correlationbased priors rather than adaptive causal structures, making them unable to accommodate evolving confounders and changing observation-action dependencies across environments. This rigidity leads to overfitting to training-specific patterns and degraded performance under distribution shifts. To address this limitation, we propose a causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes. Our model learns unified latent states that integrate vision, language, and action, while addressing spurious correlations through a dual-level intervention mechanism: at the observation level, frequency-domain perturbations simulate superficial appearance variations to enhance perceptual robustness; at the representation level, cross-episode confounder buffers perform counterfactual substitution to approximate the influence of latent confounding factors. Beyond static world modeling, our framework continuously evolves, refining these proxy representations across episodes, enabling efficient adaptation to previously unseen environments. Building on this evolving causally-inspired foundation, our world model supports counterfactual reasoning and strengthens generalization across diverse navigation contexts. Extensive evaluations on established VLN-CE benchmarks demonstrate that our method outperforms existing approaches, delivering superior navigation performance across diverse scenarios. Real-world robot evaluations further validate the practicality of our approach. Code is available in the Supplementary Material.
AWM-VLA is presented, a unified framework that embeds aligned world modeling directly inside a diffusion-transformer policy and is compatible with any diffusion or flow-matching policy, making aligned world modeling an inexpensive, broadly applicable component of generalist manipulation.
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enab...
Shaunak A. Mehta, Ananya Hazarika, Haochen Zhang et al.· Trans. Mach. Learn. Res.· 0 citations
This work proposes BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations, and introduces an asynchronous rectified-flow inference strategy wit...
Bing Zhan, Shu-Yao Shang, Shuo Lu et al.· 1 citation
AR-WAM is presented, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt and a learnable operation token dictating the atomic skill to execute, predicting scene evolution within compact latent states while decoding actions.
Yi-Cheng Jiang, Ze-Sen Gan, Xiao-Bo Wang et al.· 0 citations
The proposed hierarchical long-horizon VLA architecture with an explicit language-memory module improves the success rate and robustness of VLA models on complex long-horizon tasks while providing an interpretable semantic account of the decision process.
This work uses intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions and shows that these navigation policies are sensitive to all input modalities and do not depend on a single one.