Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, whi...
Yu Li, Sheng-He Hu, Yu-Han Wang et al.· 0 citations
PACE (Proprioception-Anchored Cross-Modal Encoder) is presented, which supervises temporal visual and F/T representations by predicting proprioceptive state transitions and is robust to perturbations that substantially degrade pose-based and learned-fusion baselines.
Yu-Han Wang, Yurou Chen, Hong-Ye Jiang et al.· 0 citations
Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-r...
Xiaomi Embodied Intelligence Team, University of Macau Shaoqing Xu, Fang Li et al.· 0 citations
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T...
Yu-Han Wang, Yurou Chen, Hong-Ye Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.