This work introduces additional 3D spatial prior information into both the training and inference stages of video diffusion models to enhance the spatial structural consistency of generated videos and introduces an energy-function guidance strategy (Warp-Guidance) driven by warping priors during denoising.
Hong-Zhou Zhu, Xue Yang, Min Zhao et al.· 0 citations
Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model, is presented and the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing is explored.
Jintao Zhang, Kai Jiang, Jin-Tao Chen et al.· 3 citations
Motus2 is presented, a self-evolving general world model for dexterous manipulation that combines egocentric data scaling and closed-loop general world model scaling to provide a general path toward self-evolving dexterous manipulation.
Hong-Zhe Bi, Zikun Zhou, Yihao Tang et al.· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.