Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can prov...
Yu-Tong Wang, Xing-Tong Ge, En-Huai Liu et al.· 0 citations
Every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions, in \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement.
Bo-Tong Zhao, Fangjie Yu, Tim Yu et al.· 0 citations
Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model, is proposed.
Haoning Yang, Xinyuan Chen, Yaohui Wang et al.· ACM Transactions on Multimed...· 1 citation
DeforM is proposed, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions, and introduces a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks.
Yunyi Li, Yu Qiao, Yaohui Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.