General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani...
Zi-Jie Diao, Yi-Tong Chen, Si-Cheng Xie et al.· 0 citations
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by...
Si-Cheng Xie, Yi-Tong Chen, Hai-Dong Cao et al.· 0 citations
Imitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizo...
Haidong Cao, Wenjun Cao, Quanhao Li et al.· 0 citations
FineVLA, an open framework for action-aligned fine-grained VLA supervision, is introduced and the largest real-world gains appear on pose, color, and approach direction, and approach direction--factors where goal-level instructions provide no guidance.