The Causal Semantic World Action Model (CSWAM) is presented, which augments FastWAM with a causal semantic expert built on V-JEPA 2.1, which provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details.
Tian-Bin Liu, Jian Zhu, Taiyi Su et al.· 0 citations
DSWAM is introduced, a Dual-System World Action Foundation Model for fine-grained robot manipulation and built and evaluated under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria.
Jian Zhu, Jianjun Zhang, Taiyi Su et al.· arXiv.org· 2 citations
MECo-WAM is proposed, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph.
Jianjun Zhang, Jian Zhu, Taiyi Su et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.