Vorch-Omni is presented, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation that supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video tr...
Vorch Team, Xiaoyu Chen, Yang Ding et al.· 0 citations
Vorch-Director is proposed, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training to produce more realistic autoregressive histories while retaining efficient teacher-forcing training.
Despite the success of diffusion models in Video Frame Interpolation (VFI), existing methods still suffer from two critical limitations. First, latent diffusion inevitably loses fine-grained details when reconstructing images from latent representations back to the pixel space. Second, multi-step sampling incurs prohib...
Zihao Zhang, Haoyu Zhao, Siqian Yang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.