Skip to content

M4World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

Jul 2026 · arXiv.org · Vol abs/2607.14005 · 1 citation · 74 references
Computer Science

TL;DR

M$^\text{4}$World, a Multi-view and Multimodal generative driving world model that synthesizes future surround-view video streams and synchronized LiDAR scans while supporting interactive object Manipulation and stable Minute-long streaming is presented.

Abstract

Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-level controllability and long-horizon stability. We present M$^\text{4}$World, a Multi-view and Multimodal generative driving world model that synthesizes future surround-view video streams and synchronized LiDAR scans while supporting interactive object Manipulation and stable Minute-long streaming. Fine-grained object manipulation is realized through a flexible conditioning interface that supports explicit control over both the spatial layout and visual appearance of individual objects. Stable minute-long streaming, on the other hand, is achieved through a multi-stage training framework that enables online causal generation in only four denoising steps while maintaining coherent world dynamics throughout extended rollouts. Building on these components, we introduce an efficient few-clip post-training as well as a suite of visual reference-conditioned generation models, preserving general generation ability while allowing rare-case customization for long-tail controllability. To assess controllability beyond realism, we further introduce an automated VLM-based judging pipeline that evaluates scene-level condition adherence, view-wise object controllability, and cross-view object consistency. Comprehensive experiments show that M$^\text{4}$World consistently delivers high generation quality, precise controllability, and stable minute-long streaming. Together with downstream long-tail augmentation and scene editing, these results demonstrate the potential of M$^\text{4}$World for controllable, scalable driving simulation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory that enables streaming scene exploration from a single input image or text prompt and shows substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploratio...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation
Preprint Sep 2026

HelloWorld: Towards Practical Applications of Generative Driving World Models

Driving world models provide a promising route toward scalable counterfactual data generation and interactive simulation beyond recorded driving logs. Realizing this potential requires a system that can generalize across diverse scenes, respond faithfully to prescribed controls, generate coherent multi-sensor observati...

Fan Lu, Han-Shi Wang, Zi-Jing Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving

Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g...

Yu Meng, Bai-Ning Zhao, Jun-Tao Wu et al. · 0 citations
Preprint Aug 2026

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).

Yuxuan Zhang, H. Xiong, Yubo Huang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Astronex-World 1.0 is presented, an open controllable video world-model foundation that provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior.

Xin Zhou, Cong-Wen Miao · 2 citations
Preprint Aug 2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Nan Duan, Hao-Yang Huang, Wei-Yang Jin et al. · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.