Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-r...
Xiaomi Embodied Intelligence Team, University of Macau Shaoqing Xu, Fang Li et al.· 0 citations
Imitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizo...
Haidong Cao, Wenjun Cao, Quanhao Li et al.· 0 citations
FM-VLA is proposed, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation, and achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches.
Ruicheng Li, Qi-Xiu Li, Ruichun Ma et al.· arXiv.org· 1 citation
This work introduces HandWorld, a unified generative framework that focuses on hand-object interaction and jointly models ego-centric videos and hand actions and learns shared cross-domain conditions through a dual-branch condition network that integrates information from both video and action domains.
Zhihao Sun, Zhi-Ying Du, Xitong Yang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.