Skip to content
Preprint

Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

A feed-forward vehicle asset reconstruction method that leverages two complementary priors to reconstruct 3D Gaussian representations for vehicles using sparse one-sided observations and significantly outperforms existing approaches in both vehicle asset completeness and geometric accuracy is proposed.

Abstract

High-fidelity vehicle assets are essential for controllable traffic scene generation, particularly for synthesizing rare and safety-critical long-tail scenarios. However, reconstructing a reusable vehicle representation from in-the-wild onboard images remains challenging for two reasons. First, image-to-3D generation methods generally produce models without reliable metric scale. Second, onboard cameras usually observe only one side of a target vehicle, making conventional multi-view reconstruction incomplete on unobserved regions. To solve these problems, we propose a feed-forward vehicle asset reconstruction method, which leverages two complementary priors to reconstruct 3D Gaussian representations for vehicles using sparse one-sided observations. To achieve metric-scale reconstruction, a visual foundation model is first utilized to serve as a visual prior for Gaussian initialization. The Gaussian attributes are then estimated by a learnable encoder-decoder module. A symmetry-aware cloning strategy is presented to complete the unobserved side directly in Gaussian space, which exploits the bilateral structure of vehicles as a geometric prior. Experiments on the public dataset demonstrate that the proposed method significantly outperforms existing approaches in both vehicle asset completeness and geometric accuracy.

View source

Similar papers

Preprint Sep 2026

RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction

Recent developments in feed-forward 3D reconstruction resulted in models which can recover dense scene representations and camera motion solely from an image stream. However, such predictions are prone to becoming inconsistent over long trajectories, specifically in demanding environments with repetitive structures, we...

Ting-Jun Huang, Dmitry Rudshin, Mathieu Meyer et al. · 1 citation
Preprint Aug 2026

USR-Drive: Unified Driving Scene Representation via Joint Denoising of 3D Gaussians and Boxes

Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typ...

Li-Heng Chen, Haokai Pang, Cheng Su et al. · 0 citations
Preprint Sep 2026

RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that combine reconstruction and generation. However, existing methods bridge the two with rendered images or explicit 3D representations such as...

Xiaolei Lang, Ze Kang, Ze-Hao Huang et al. · 0 citations
Preprint Sep 2026

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied...

Ao-Xiang Fan, Corentin Dumery, Nicolas Talabot et al. · 0 citations
Preprint Sep 2026

VDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene Reconstruction

Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In recent years, 3D Gaussian Splatting (3DGS) has demonstrated remarkable advantages in both visual quality and computational efficiency, making it a promising representation f...

Hao-Lin Yu, Jia-Dong Tang, Yi-Xian Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Sparse auto-regressive modeling for scene generation from multi-view images

Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generati...

Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.