Skip to content

Predictive Structure Improves Video Diffusion Dynamics

· 0 citations · 65 references

TL;DR

Experiments show that LDO substantially improves physical commonsense, object permanence, and trajectory fidelity while preserving visual quality, suggesting that predictive latent supervision offers a practical route to make video generators not only photorealistic but also physically legible.

View source

Similar papers

Open access Aug 2026

Long-Horizon Video Generation with Temporally Consistent Diffusion and Scene-Graph Guidance

A novel framework that integrates temporally consistent diffusion models with dynamic scene-graph guidance that structurally constrains the generative process, ensuring that objects, their attributes, and their interrelationships remain stable over extended durations is introduced.

Jacob A. Jenkins · 0 citations
Preprint Jul 2026

Gen4U: Unifying Video Generation and Understanding via Diffusion

Gen4U (Generation for Understanding), a framework that repurposes these generative representations with a single forward pass, is introduced, achieving strong perception performance while fully preserving the model's ability to generate high-quality video.

Michael King, Aravindh Mahendran, M. Grimes et al. · 0 citations
Preprint Jul 2026

Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency

This work proposes Cycle-World, a novel framework designed for stable and temporally consistent long-video generation that tackles error drift by enforcing strict temporal reversibility across both the training and inference phases, and demonstrates that forward generative drift can be strictly bottlenecked by a cycle-consistency objective.

Zihan Su, Teng Hu, Jiangning Zhang et al. · 1 citation
Preprint Jul 2026

Self-Supervised Learning of Structured Dynamics from Videos

The Structured Dynamics Model (SDM) is proposed, which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens.

Lukas Knobel, Andrew Zisserman, Yuki M. Asano · 0 citations