MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair, shows consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.
Abstract
Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.
SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.
Zongchuang Zhao, Xin Zhou, Tianyang Xu et al.· 4 citations
SV-WAM is proposed, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference and introduces a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning saf...
Jin-Yang Wang, Shi-Wei Li, Jun-Jian Wang et al.· 0 citations
World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficul...
Zhao-Xin Fan, Tian-Bao Zhang, Wen-Jun Wu et al.· 0 citations
DA-WAM is proposed, a framework that unifies predictive representation learning, action-conditioned future modeling, and trajectory scoring under a single decision-making objective and demonstrates state-of-the-art performance on NAVSIM-v1 and NAVSIM-v2.
Rui-Guo Zhong, Ben-Shan Ma, Xiao-Long Chen et al.· 3 citations
PhysWAM is presented, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer and Coupled Point Projection is introduced, demonstrating that the geometric relationship between scene depth and ego motion provides a direc...
Dhruv Parikh, Feng-Cheng Yu, Quan-Kai Gao et al.· 0 citations
Method is presented, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement and shows that it naturally supports edge--cloud deployment with substantially lower communication overhead than the b...
Yi-Xin Zheng, Jiangran Lyu, Yun-Tian Deng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.