Skip to content
Preprint

VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

Sep 2026 · 1 citation · 54 references
Computer Science

TL;DR

VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings, is introduced, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

Abstract

3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through ligh...

Jerrin Bright, John S. Zelek · 0 citations
Preprint Aug 2026

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.

Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al. · 1 citation
Preprint Sep 2026

SkyAnchor: Updating Metric-scale Aerial 3D Gaussian Scenes from Unposed Ground-View Sequences

We study how to update a pre-built aerial scene with a newly captured, unposed ground-view sequence. The aerial scene already contains a reliable metric Structure-from-Motion (SfM) reconstruction and a pre-trained 3D Gaussian Splatting (3DGS) model, whereas the ground-view sequence is collected later to add street-leve...

Zhuo-Xiao Li, Xin-Yi Liu, Tao-Yu Wu et al. · 0 citations
Preprint Sep 2026

RIGOR: Rig-Informed Geometry for Omnidirectional Reconstruction

Recent developments in feed-forward 3D reconstruction resulted in models which can recover dense scene representations and camera motion solely from an image stream. However, such predictions are prone to becoming inconsistent over long trajectories, specifically in demanding environments with repetitive structures, we...

Ting-Jun Huang, Dmitry Rudshin, Mathieu Meyer et al. · 1 citation
Preprint Sep 2026

Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge

Transferring the rich priors of large 2D foundation models to sparse 3D LiDAR remains challenging, as training native 3D foundation models at comparable scale is limited by data and annotation scarcity. We introduce a LiDAR-conditioned diffusion model trained on pseudo-labels from off-the-shelf 2D foundation models. Th...

Samed Doğan, Nico Leuze, Alfred Schöttl · 0 citations
Preprint Aug 2026

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed durin...

Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.