While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, th...
Hao-Ran Qin, Ren-Rong Wu, Tian-Yu Huang et al.· 0 citations
AViTS models spatial importance via latent-text attention and temporal importance via token-level feature variation across diffusion timesteps, and fuses them to enable spatiotemporal importance-aware selective upsampling: it prioritizes resolution refinement for critical tokens while deferring less important ones, the...
Hao-Ran Qin, Zhen Yan, Shikang Zheng et al.· 0 citations
LinCa decomposes cached features into sub-components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component, forming a unified Decompose-Predict-Reconstruct pipeline.
Jin-Shan Liu, Hao-Ran Qin, Xiao-Bing Tu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.