Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but such training relies on...
Ze Chen, Pei-Dong Liu, Jiawei Li et al.· 0 citations
WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM), is introduced.
Pei-Dong Liu, Zhi-Yuan Xiang, Ming-Yang Li et al.· 0 citations
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understandi...
Hanwen Wan, Dafeng Chi, Lin-Bo Zhai et al.· 0 citations
This work introduces Event3R, a feed-forward framework that directly maps asynchronous event streams to globally consistent 3D point clouds, and proposes a Masked Bin Modeling strategy for self-supervised pre-training, enabling robust temporal representation learning with minimal labeled data.