Recent diffusion and autoregressive models have substantially improved text-to-video generation, yet producing coherent long-form story videos with consistent characters and aligned audio remains challenging. Existing methods often suffer from narrative drift, unstable character identity, weak cross-scene continuity, a...
Ting Huang, Biao Wu, Rong-Hao Chen et al.· 0 citations
The conditional generation of 3D motion has emerged as a key research topic due to its wide applicability across robotics, AR/VR, gaming, and content creation. However, extending recent advances in text-driven human motion generation to the animal domain remains challenging due to two core limitations. First, animals e...
Ze-Yu Zhang, Zhi-Yuan Zhang, Siheng Wang et al.· 0 citations
This work proposes FlashMo, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design, and introduces MotionSiT, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over SO(3) manifolds, enab...
Zeyu Zhang, Yiran Wang, Danning Li et al.· Neural Information Processin...· 11 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.