Experiments show that LDO substantially improves physical commonsense, object permanence, and trajectory fidelity while preserving visual quality, suggesting that predictive latent supervision offers a practical route to make video generators not only photorealistic but also physically legible.
This work introduces Latent Dynamics Reasoning (LDR), the first video world model that extrapolates learned dynamics beyond its training distribution and can even generalize under severe shift.
Haodong Li, Shaoteng Liu, Tianyu Wang et al.· 0 citations
A novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance that structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations is introduced.
Jacob A. Jenkins· Journal of innovative resear...· 0 citations
Gen4U (Generation for Understanding), a framework that repurposes these generative representations with a single forward pass, is introduced, achieving strong perception performance while fully preserving the model's ability to generate high-quality video.
Michael King, Aravindh Mahendran, M. Grimes et al.· 0 citations
This work proposes Cycle-World, a novel framework designed for stable and temporally consistent long-video generation that tackles error drift by enforcing strict temporal reversibility across both the training and inference phases, and demonstrates that forward generative drift can be strictly bottlenecked by a cycle-consistency objective.
Zihan Su, Teng Hu, Jiangning Zhang et al.· 1 citation
The Structured Dynamics Model (SDM) is proposed, which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens.
Lukas Knobel, Andrew Zisserman, Yuki M. Asano· 0 citations