Skip to content
Preprint

World in World: Explore the World with World Models

Sep 2026 · 0 citations · 79 references
Computer Science

TL;DR

World in World is presented, a training-free inference-time interface that converts heterogeneous control evidence into camera- and time-labelled clean visual states, which are read through the native self attention of a frozen causal video model.

Abstract

Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded event, place observed content in the requested view, plausibly complete newly exposed regions, and recover previously generated appearance on revisits. Existing methods typically address these requirements through task-specific modules or additional training. We present World in World, a training-free inference-time interface that converts heterogeneous control evidence into camera- and time-labelled clean visual states, which are read through the native self attention of a frozen causal video model. The evidence comprises source-video observations, target-view scene projections, geometry renderings that guide completion of newly exposed subject regions, and retrieved generated states beyond the rolling cache. Each evidence source carries token-level support and its own availability schedule. A correspondence router combines persistent point identities with geometry to establish token correspondences, guiding supported queries towards matching source-video tokens. Evidence-wise attention CFG (EWA) then independently regulates each auxiliary channel's additional contribution using attention responses from the same denoising forward pass. The shared interface supports camera-controlled rerendering, long-horizon revisiting, and human-motion transfer with the same frozen backbone. We evaluate World in World on camera-controlled video rerendering under diverse viewpoint changes, assessing perceptual quality, temporal consistency, and camera-following accuracy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory that enables streaming scene exploration from a single input image or text prompt and shows substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploratio...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation
Preprint Aug 2026

Sekai2: From World Exploration to Interactive World Modeling

Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant tem...

Kang He, Wen-Shuo Peng, Zi-Hui Gao et al. · 2 citations · ⚡1
Preprint Aug 2026

Can Video World Models Track Unobserved World States?

Video world models are increasingly used as simulators, but visual fidelity alone does not show that a model maintains the hidden state of the world. We examine this difference with an action-conditioned video Shell Game, a visual analogue of $S_5$ state tracking that separates visual rendering from compositing the uno...

Joonghyuk Shin, Yicong Hong, Jaesik Park et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Astronex-World 1.0 is presented, an open controllable video world-model foundation that provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior.

Xin Zhou, Cong-Wen Miao · 2 citations
Preprint Oct 2026

ActiveWAM: Evidence-Aware Active Vision for World-Action Models

Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate...

Ren-Jun Wu, Lu-Zhou Ge, Xue-Song Li · 0 citations
Review Sep 2026

The Past Frames the Future: Memory for Autoregressive Video Generation

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands,...

Harold Haodong Chen, Rong-Jin Guo, Di-Sen Lan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.