World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined futures with limited adaptation to evolving execution progress, potentially introducing distracting or unreliable predictive cues. This limi...
Yi-Jie Zhu, Zi-Tong Yu, Wei Li et al.· 0 citations
This approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model, achieving a speedup of 1.8$\times to 2.5$\times compared to the baseline VLLM.
Zihan Song, Shuo Ye, Bo Zhao et al.· arXiv.org· 0 citations
This work proposes GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling, and introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions.
Taorui Wang, Wei Xia, Hui Ma et al.· arXiv.org· 0 citations
AC-VLA is introduced, a plug-and-play Action Compositional learning framework comprising two architecture-agnostic components that achieves a ~28% absolute improvement on compositional OOD tasks while maintaining near-perfect in-distribution performance.
Xiaojiang Peng, Kai Peng, Jie Lu et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.