Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations...
Xiao-Da Yang, Can Wang, Yu-Xiang Liu et al.· 0 citations
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whos...
AgiBot Research Team, Renhang Liu, Wen-Zhi Zhao et al.· 2 citations· ⚡1
Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
Xiaowei Cai, Yunuo Cai, Bing Chen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.