Skip to content

GeoWorldAD: Geometry World Action Model for Autonomous Driving

Jul 2026 · arXiv.org · Vol abs/2607.17521 · 1 citation · 80 references
Computer Science

TL;DR

GeoWorldAD is proposed, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens and progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement.

Abstract

Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.

View source

Similar papers

Preprint Aug 2026

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-tr...

Yiren Lu, Xin Ye, Jiaming Liu et al. · 1 citation · ⚡1
Preprint Sep 2026

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric...

Dhruv Parikh, Feng-Cheng Yu, Quan-Kai Gao et al. · 0 citations
Preprint Aug 2026

UniNav: A Unified World-Action Diffusion Model for Visual Navigation

UniNav is presented, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process, unifying future prediction and action generation in a shared framework and introduces two variants: UniNav-Full and UniNav-Fast.

Changqing Zhou, Yueru Luo, Ze-Yu Jiang et al. · 2 citations
Preprint Sep 2026

DroneWAM: Efficient World Action Model for Drone Visual Navigation

World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone...

Liang Yao, Fan Liu, Hong-Bo Lu et al. · 0 citations
Preprint Sep 2026

READ: Learning Risk-Informed Fields for End-to-End Autonomous Driving

Autonomous driving requires more than recognizing what is present in a scene: a planner must determine how road structure, surrounding agents, and their motion states should influence a future maneuver. Existing learning-based planners can capture these influences through latent scene features and trajectory decoders,...

Zhi-Yuan Liu, Yuan-Xin Tian, Ze-Hong Ke et al. · 0 citations
Preprint Aug 2026

Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

Geo-VLA is proposed, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations and introduces Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning.

Ran Chen, Jiaxing Ren, Zhikun Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.