SV-WAM is proposed, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference and introduces a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness.
Abstract
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.
Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual struc...
Yi-Guang Yang, Jian-Kun Peng, Xiao-Ming Wang et al.· 0 citations
MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair, shows consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.
Shuai Liu, Hechangle Gong, Hao Jiang et al.· 0 citations
Vid2WAM is proposed, an offline distillation framework that transfers visual diffusion priors from a large video foundation model into a compact WAM student and introduces source-aware residual action adaptation that learns source-specific corrections around a shared action backbone and mitigates interference from nois...
Chen-Hao Qiu, Ruixiang Wang, Runyi Zhao et al.· 2 citations
Vehicle-infrastructure cooperation can complement onboard sensing with broader and more informative observations of the traffic environment, providing valuable support for end-to-end autonomous driving. However, existing cooperative driving methods mainly exploit roadside information to enhance the representation of th...
Jun-Wei You, Wei-Zhe Tang, Can Wang et al.· 0 citations
World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-level task evolution,...
Jie Wu, Yu-Zhi Huang, Jun-Qi Liu et al.· 0 citations
Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate...
Ren-Jun Wu, Lu-Zhou Ge, Xue-Song Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.