Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can prov...
Yu-Tong Wang, Xing-Tong Ge, En-Huai Liu et al.· 0 citations
Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) prim...
Fang-Yu Lin, Xing-Tong Ge, Lu Zhu et al.· 0 citations
Omni-LiveAvatar is presented, the first framework for minute-level, real-time streaming joint audio-video avatar generation and proposes a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary st...
Lu Zhu, Xing-Tong Ge, Fang-Yu Lin et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.