Skip to content

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

The Causal Semantic World Action Model (CSWAM) is presented, which augments FastWAM with a causal semantic expert built on V-JEPA 2.1, which provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details.

Abstract

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.

View source

Similar papers

Preprint Sep 2026

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM), is introduced.

Pei-Dong Liu, Zhi-Yuan Xiang, Ming-Yang Li et al. · 0 citations
Preprint Sep 2026

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embo...

Hao Wang, Jia-Jun Wen, Jing-Zhi Liu et al. · 0 citations
Preprint Oct 2026

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...

Jia-Yi Chen, Wen-Xuan Song, Jing-Bo Wang et al. · 0 citations
Preprint Aug 2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...

Yi-Han Lin, Jia-Wei He, Shi-Feng Bao et al. · 10 citations · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.