Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a wor...
Ze-Tao Cai, Ya-Ping Li, Yi-Qun Wang et al.· 0 citations
The Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects, is introduced, a structured intermediate representation that provides a strong inductive bias aligned with the underlying structure of articulated motion.
Ya-Ping Li, Zhaxizhuoma, Qiao-Jun Yu et al.· arXiv.org· 0 citations
This work revisits the scaling recipe for BFMs and demonstrates that substantial performance gains can be achieved through the coordination of three core components: the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the...