While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, an...
Guiyu Zhao, Long-Teng Guo, Yang-Hong Mei et al.· 1 citation
World Tokens is an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation and is highly competitive on LIBERO, attains the best reported averages on SIMPLER, and substantially improves real-world R1 Pro success over a matched...
Qu Tang, Benhui Zhuang, Bo Yuan et al.· 1 citation· ⚡1
TimeThink is proposed, a reinforcement learning framework that explicitly guides temporal evidence discovery in Video-LLMs and introduces a step-wise temporal process reward that provides localized credit assignment for these clues and a joint process--outcome optimization objective that balances reasoning fidelity wit...
Handong Li, Longteng Guo, Zikang Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.