Skip to content

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

Jun 2026 · arXiv.org · Vol abs/2606.27964 · 1 citation · 43 references
Computer Science

TL;DR

This work proposes "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control, and introduces a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence.

Abstract

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accumulation and temporal degradation during extended rollouts. This issue is further amplified under heterogeneous controls such as human motion and camera trajectories, which may interfere and destabilize a pretrained video prior, while existing methods often trade off controllability and visual quality. We propose"Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control. Our key idea is to decouple control learning while preserving a unified autoregressive video prior. We introduce a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence. For human motion control, we design a t-guided Dynamic Projection mechanism and a refined Motion-CFG strategy, enabling temporally smooth and accurate motion alignment without degrading visual fidelity, and supporting multi-person control.After learning a robust motion prior, we introduce a second-stage camera-trajectory control module to compose human dynamics with viewpoint changes for coherent world exploration. We further construct a large-scale dataset with synchronized video, text, human-motion, and camera-trajectory annotations, organized into motion-centric and camera-centric subsets for decoupled training. Extensive experiments show stable long-horizon generation with precise controllability and high visual quality. See more at https://whydahuzi.github.io/Directing-the-World.github.io/.

View source

Similar papers

Preprint Jul 2026

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk.

AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al. · 2 citations
Open access Jul 2026

Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

ARDY is introduced, a streaming generation framework that bridges this gap by enabling high-fidelity motion generation controllable via online text prompts and flexible kinematic constraints and balancing precise trajectory control with efficient generative learning.

Kaifeng Zhao, Mathis Petrovich, Haotian Zhang et al. · 2 citations
Preprint Aug 2026

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

MotionCraft is presented, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface to deliver temporally consistent, high-quality reconstructions under streaming constraints.

Rong Fu, Chunlei Meng, Yangcheng Zeng et al. · 0 citations
Preprint Jul 2026

Mitigating Compounding Error via Video Representation Regularization

This work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.

Taiye Chen, Qi Zhang, Yisen Wang · 0 citations
Jun 2026

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Walking in the Implicit, a scene-centric paradigm that changes the rollout variable from frame latents to a fixed-length, renderable implicit state, termed Neural Implicit Scene (NIS), factorizes interactive generation into stochastic transition of a compact scene state and deterministic pose-conditioned rendering given the sampled state.

Zhiqi Li, Chengrui Dong, Zhenhua Du et al. · 0 citations
Preprint Jul 2026

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure and consistently improves temporal quality across multiple autoregressive diffusion models.

Thanh-Nhan Vo, Trong-Thuan Nguyen, T. Le et al. · 0 citations