We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation o...
Yi-Yi Cai, Yu-Han Wu, Kun-Hang Li et al.· 0 citations
Real-world capture is heterogeneous: perspective, fisheye, and $360^\circ$ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are either informed of th...
Qiao-Ge Li, Yi-Fan Zhan, Hai-Jun Yang et al.· 0 citations
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requi...
Chen-Chieh Liao, Yi-Chen Peng, Yi-Yi Cai et al.· 0 citations
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as...
Fang-Yuan Tu, Xiang-Yue Zhang, Yi-Yi Cai et al.· 0 citations
Triangular Resampling is introduced, a post-training method for mitigating long-horizon error accumulation in motion diffusion models that addresses the mismatch between ground-truth-derived training windows and model-generated inference states.
Kun-Hang Li, Yi-Yi Cai, Xiang-Yue Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.