Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the generator's evolving comp...
Wei-Li Zeng, Feng Tian, Sheng-Qi Liu et al.· 0 citations
This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.
Wei-Li Zeng, Yi-Tong Xing, Fu-Long Liu et al.· 1 citation
A unified spatiotemporally decoupled framework named DeMoDiff is proposed, which jointly redesigns representation and architecture and incorporates spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability.
Cheng-Qun Yang, Liang Xu, Yan-Ping Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.