Skip to content
Preprint

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

Aug 2026 · 1 citation · 99 references
Computer Science

TL;DR

ABot-Recon is presented, a simple streaming model that caches KV features from only the preceding 11 frames and predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose that remains equivariant under changes of reference frame.

Abstract

Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of $0.12^\circ$, reducing both errors by approximately 40\% relative to the best prior results.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

This work proposes LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency, and introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows w...

Jing-Ke Zhou, Chen-Hang Ma, Zhi-Zhou Zhong et al. · 1 citation
#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory that enables streaming scene exploration from a single input image or text prompt and shows substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploratio...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation
Preprint Aug 2026

GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly

GeoWeaver, a unified framework comprising a Geometric Prior Model (GPM) and Test-Time Adaptation (TTA) and a robust CDF-style objective jointly optimizes weighted 2D reprojection and 3D consistency residuals, is presented.

Tinghao Jiang, Sheng Tang, Shengzhe Wei et al. · 0 citations
Preprint Aug 2026

UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate super...

Tiancheng Chen, Sheng Tang, Wenhua Jin et al. · 1 citation
Preprint Aug 2026

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but dire...

K. M. Le, H. Pham, Danh Thanh Luu et al. · 1 citation
Preprint Sep 2026

Seeing the World and the Self from Egocentric Video

Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.

Kai Guan, Minchao Jiang, Ruichen WangLi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.