Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action...
Yi-Ming Jiang, Jin Chen, Chong-Yang Xu et al.· 0 citations
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid...
Jin Chen, Yi-Ming Jiang, Chong-Yang Xu et al.· 0 citations
This work presents RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling, and introduces RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states.
Ziqin Wang, Hao Li, Weijun Wang et al.· arXiv.org· 0 citations
StageWAM is introduced, which augments a Motus-based World Action Model with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor, which uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage.
Xiao Liu, Yuguang Yang, Xi Wang et al.· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.