AffordanceWAM is introduced, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World, and supports affordance as an effective interface for both vision-language-action learning and huma...
Jia-Di You, Qi-Ze Yu, Yue Chen et al.· 0 citations
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction d...
Jun-Feng Li, Junjie He, Zhi-De Zhong et al.· 1 citation
The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.
Hao-Dong Yan, Jun-Feng Li, Jun-Jie He et al.· 2 citations
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.