Skip to content
Preprint

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Aug 2026 · 0 citations · 47 references
Computer Science

TL;DR

VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting and achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.

Abstract

Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex decoders or auxiliary cues, they remain bottlenecked by the weak geometric capacity of upstream features. We argue that leveraging pretrained visual geometry priors strengthens upstream representations and alleviates the geometric ambiguity in sparse surround views. To this end, we propose VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting. First, VGGD leverages VGGT to provide transferable multi-view geometric prior tokens. Next, we introduce a Dual-Path Neck to decouple geometry-consistent and appearance-aware representations, improving appearance completion in weakly observed regions. We further apply Scale Warmup to stabilize early geometry learning and suppress scale drift under ego-pose changes. Finally, we use a hybrid pixel--volume Gaussian decoder to produce a renderable 3D Gaussian scene for novel-view synthesis. Experiments on the nuScenes single-frame benchmark show that VGGD achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

This work presents VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis, and introduces robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness.

Kang-Jie Chen, Xiang-Yu Li, Dong-Bin Zhang et al. · 0 citations
Preprint Aug 2026

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

UniWorld-View is introduced, a unified framework for controllable large-baseline novel view synthesis from monocular inputs that integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation.

Hai-Yang Zhou, Wang-Bo Yu, Chaoran Feng et al. · 2 citations
Preprint Sep 2026

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation

Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-view surround camera rigs provide broad scene coverage, but the spatially adjacent images typically overlap only minimally. Consequently, the depth of most pixels must be inferred from monocular appearance cues....

S. Abualhanud, M. Mehltretter · 0 citations
Preprint Aug 2026

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.

Xinhui Liu, Can Wang, Wei Jiang et al. · 1 citation
Preprint Aug 2026

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.

Yu-Fei Zhang, Chen-Lu Zhan, Hong-Wei Wang · 0 citations
Preprint Aug 2026

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

GeoFlow is a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors, using a Geometry-Aligned Prior (GAP) distribution as starting point, and can achieve remarkable efficiency of both training and inference.

Jia-Zhen Liu, Hangbiao Li, J. Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.