Jun 2026· arXiv.org· Vol abs/2606.27964· 1 citation· 43 references
Computer Science
TL;DR
This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.
Abstract
Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accumulation and temporal degradation during extended rollouts. This issue is further amplified under heterogeneous controls such as human motion and camera trajectories, which may interfere and destabilize a pretrained video prior, while existing methods often trade off controllability and visual quality. We propose"Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control. Our key idea is to decouple control learning while preserving a unified autoregressive video prior. We introduce a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence. For human motion control, we design a t-guided Dynamic Projection mechanism and a refined Motion-CFG strategy, enabling temporally smooth and accurate motion alignment without degrading visual fidelity, and supporting multi-person control.After learning a robust motion prior, we introduce a second-stage camera-trajectory control module to compose human dynamics with viewpoint changes for coherent world exploration. We further construct a large-scale dataset with synchronized video, text, human-motion, and camera-trajectory annotations, organized into motion-centric and camera-centric subsets for decoupled training. Extensive experiments show stable long-horizon generation with precise controllability and high visual quality. See more at https://whydahuzi.github.io/Directing-the-World.github.io/.
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· 2 citations
ARDY is introduced, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints and balancing precise trajectory control with efficient generative learning.
Kaifeng Zhao, Mathis Petrovich, Haotian Zhang et al.· ACM Transactions on Graphics· 2 citations
MotionCraft is presented, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface to deliver temporally consistent, high-quality reconstructions under streaming constraints.
Rong Fu, Chunlei Meng, Yangcheng Zeng et al.· 0 citations
This work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.
Walking in the Implicit, a scene-centric paradigm that changes the rollout variable from frame latents to a fixed-length, renderable implicit state, termed Neural Implicit Scene (NIS), factorizes interactive generation into stochastic transition of a compact scene state and deterministic pose-conditioned rendering given the sampled state.
Zhiqi Li, Chengrui Dong, Zhenhua Du et al.· arXiv.org· 0 citations
SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure and consistently improves temporal quality across multiple autoregressive diffusion models.
Thanh-Nhan Vo, Trong-Thuan Nguyen, T. Le et al.· 0 citations