Skip to content
Preprint

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

Sep 2026 · 0 citations · 67 references
Computer Science

TL;DR

WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM), is introduced.

Abstract

Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $\pi_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.

View source

Similar papers

Preprint Sep 2026

From World Models to World Action Models: Rethinking Next-State Prediction

Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets,...

Ting-Yu Yuan, Zi-Ming Ji, Biao-Liang Guan et al. · 1 citation
Preprint Aug 2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy, which matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

Jing-Kai Wang, Zihan Tang, Gu Zhang et al. · 0 citations
Preprint Aug 2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...

Yi-Han Lin, Jia-Wei He, Shi-Feng Bao et al. · 10 citations · ⚡1
Preprint Aug 2026

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

LD4WAM is presented, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated fut...

Zhen Shen, Jia-Qi Liang, Jasper Lu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

The Causal Semantic World Action Model (CSWAM) is presented, which augments FastWAM with a causal semantic expert built on V-JEPA 2.1, which provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details.

Tian-Bin Liu, Jian Zhu, Taiyi Su et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.