Preprint
Jul 2026
Video Generation Models are General-Purpose Vision Learners
GenCeption is introduced, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions, and suggests that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world.
Letian Wang, Chuhan Zhang, Rishabh Kabra et al.
· 6 citations