Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pixel-, and latent-spac...
Jun-Hao Xiao, Hao-Xiang Zhao, Meng-Hao Fang et al.· 0 citations
G0.5 is introduced, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective, which exceeds state-of-the-art models across 7 independent regimes.
Yi-Cheng Liu, Zibin Dong, Baijun Ye et al.· 16 citations· ⚡2
MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization.
Ze-Hua Fan, Jun-Jie He, Wen-Xuan Song et al.· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.