Action-conditioned cloth dynamics prediction requires both locally plausible deformation and long-range coordination. Existing approaches largely follow two paradigms. Mesh-based GNNs capture local physical responses through material connectivity. However, their finite message-passing range limits coordination between...
Zi-Hang Wang, Jian-Ming Hu, Shang Su et al.· 0 citations
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce...
Ze Feng, Yi-Xu Feng, Ling-Yu Xiao et al.· 0 citations
This work proposes a self-evolving method that reduces failure rates by 51--67% relative to trained baselines and by 8-25% relative to state-of-the-art vision-language-action models after replacing redundant nominal scenarios with diverse failure-prone ones.
Linxuan He, Yuying Tian, Ling-Xiang Fan et al.· arXiv.org· 0 citations
This work proposes DC-WAM, a dynamic-centric WAM framework that redistributes supervision and computation in the RGB video branch that consistently improves policy performance, especially under out-of-distribution perturbations in lighting, object appearance, and background texture.
Haoyuan Ji, Lingxiang Fan, Shang Su et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.