World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficul...
Zhao-Xin Fan, Tian-Bao Zhang, Wen-Jun Wu et al.· 0 citations
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to localize try-on regions, making them vulnerable to large motions and severe occlusion...
Wei Zhang, Xin Li, Pei-Shu Shi et al.· 0 citations
H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model, and introduces temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions.
Dan Chen, Ze-Qing Wang, Zibin Lin et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.