With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source clip frame by frame over a fixed time span. This pattern fails for open-ended streams, e.g., restyling...
Yunze Tong, Mu-Shui Liu, Can-Yu Zhao et al.· 0 citations
Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.
Shi-Yi Zhang, Mu-Shui Liu, Yunze Tong et al.· 0 citations
VLTok is a novel 1D hybrid tokenizer that unifies V isual and L anguage representations in a shared Tok en space through a self-prompted training paradigm, and achieves state-of-the-art performance in both image reconstruction and image generation.
Hualiang Wang, Siming Fu, Wei-Nan Jia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.