Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views, and real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
This work presents Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance, and proposes an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8....
Jia-Min Zhou, Qihang Zhang, Gangwei Xu et al.· 8 citations
TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.