World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, o...
Bingqi Huang, Bingchuan Wei, Ying-Kai Cai et al.· 2 citations
Permanent lunar habitation will require robotic systems that can maintain infrastructure, recover from local failures, and acquire new operational capabilities under limited Earth supervision. Existing planetary robots are largely fixed-function specialists, while end-to-end foundation-model policies remain difficult t...
Bingqi Huang, Bingchuan Wei, Ying-Kai Cai et al.· Astronautics· 0 citations
Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, de...
Bingqi Huang, Bingchuan Wei, Xuan Wang et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.