Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas...
Ke-Rui Ren, Tao Lu, Lin-Ning Xu et al.· 0 citations
Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a wor...
Ze-Tao Cai, Ya-Ping Li, Yi-Qun Wang et al.· 0 citations
An efficient generative framework designed to improve in-the-wild robustness under diverse real capture conditions and demonstrate strong perceptual quality, semantic fidelity, and temporal consistency on unseen videos, as well as improved robustness in downstream 3D reconstruction under severe motion blur.