Skip to content
Open access

Using textureless, low-detailed 3D city models for visual localization

Aug 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · 0 citations · 31 references

TL;DR

This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.

Abstract

Abstract. Accurate camera pose estimation in urban environments remains challenging when reference imagery is generated from low-detailed, textureless 3D city models and must be matched against real world imagery. In this work we (i) extend our existing iterative object-basesd visual localization approach with an additional semantic feature and (ii) conduct a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs. As a first step to close the domain gap, we augment our iterative object-based visual localization pipeline with semantic masks derived from a pretrained semantic segmentation model. Intersection-over-Union between query and rendered masks is incorporated into the matching score, leading to a better pose accuracy. For the baseline study, we use a range of feature matching techniques: handcrafted (SIFT, AKAZE, ORB, FAST), learned detectors (XFeat, Key.Net, DeDoDe, DISK, AffNet), learned descriptors (XFeat, DISK, DeDoDe, HardNet), learned matchers (LightGlue, LoFTR), the line matcher SOLD2, and the learned matchers MINIMA-RoMa, MINIMA-LoFTR, MINIMA-XoFTR, and MatchAnything, which were trained on cross-modality datasets. The cross-modality focused matchers achieved the best results. For 20% 10%, 9%, and 7% of the evaluated query images the estimated camera pose had a translation error less than 5m and a rotation error less than 5◦. In this context, the other methods were only able to achieve a maximum success rate of 1.4%.

Read PDF

Similar papers

Preprint Jul 2026

Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis

This work proposes a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS), and shows substantial gains in relocalization accuracy under challenging conditions.

M. Peribañez, Javier Civera, Rudolph Triebel et al. · 0 citations
2026

Revisiting Visual Localization: A Feed-Forward Localization Framework With a Lightweight Scene Representation

Visual localization is a key technology in many vision-based measurement applications, aiming to estimate the camera pose of a query image in a known environment. However, most existing methods rely on heavy scene-specific representations, such as explicit 3-D map construction or per-scene training. Constructing and maintaining such representations introduces nonnegligible computational overhead, storage burden, and long-term maintenance costs. To address this issue, we propose a novel visual localization pipeline that uses a set of posed reference images as a lightweight scene representation and localizes query images without explicit 3-D map construction or scene-specific training. Specifically, we exploit a geometric foundation model to infer local multiview geometry from the query image and its retrieved references. Since the predicted geometry is expressed in an arbitrary local coordinate system with unknown scale, a key challenge is how to recover an accurate metric pose of the query image from such local predictions. To address this challenge, we design a global pose recovery strategy that first registers the predicted local geometry to the world coordinate system through joint center–orientation similarity alignment using the posed reference images as global anchors, and then refines the query pose by optimizing query-associated 3-D landmarks under multiview 2-D–3-D geometric constraints. The experimental results on multiple benchmark datasets show that our method achieves competitive localization performance and improved robustness under sparse reference-view settings and challenging viewpoint or appearance variations, reducing the average translation and rotation errors of the strongest Unseen baseline from 55 cm and 0.56° to 13 cm and 0.23° on Cambridge Landmarks, respectively.

Wenhao Lin, Cong Guo, Yu Wu et al. · 0 citations
Preprint Jul 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

6D pose estimation remains a key challenge in robotics and computer vision, particularly in industrial environments. The deployment of currently available data-driven methods is often limited by resource-intensive data pipelines, reliance on textured 3D models, and sensitivity to geometric deviations caused by damages or assembly defects. We present PIXIE, a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model. Synthetic depth and normal maps are rendered from sampled reference viewpoints and matched to the query image via a pretrained cross-modality feature matcher. Matched keypoints are back-projected to obtain 2D--3D correspondences for PnP-based pose estimation. Relying exclusively on geometry makes the method inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object. We evaluate on widely-used public benchmarks, reporting state-of-the-art results on texture-less objects without object-specific training, and introduce a novel dataset with assembly defects, texture variations, and occlusion to demonstrate real-world applicability.

Leon Jungemeyer, A. Magaña, Gautham Mohan et al. · 0 citations
Preprint Jul 2026

MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors

Classical image correspondence is solved at the level of sparse keypoints or dense pixels, but the systems that consume these matches - object-level mapping, topological navigation, scene-graph maintenance - reason about whole objects. Recent work narrows this gap by matchng directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are obtained by pooling features from large 3D foundation models over the masks. We build on this segment-level matching paradigm and propose three learned matching heads: a LightGlue-style attention head with DoubleSoftmax scoring on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from the VGGT foundation model before pooling; and - as our main contribution - a multi-view extension that performs joint self-attention over segments drawn from several views at once, recovering transitive correspondences that strictly pairwise matchers cannot reach. Under a stratified zero-shot protocol on Replica and Virtual KITTI 2 with controlled viewpoint baselines from 0 deg to 180 deg, the LightGlue-style head improves over a parameter-free Sinkhorn matcher on the same MASt3R backbone by +4.85 AUPRC on Replica and +25.9 AUPRC on Virtual KITTI 2. Dropped into the RoboHop topological navigation pipeline on the Habitat-Matterport 3D (HM3D) Instance Image Navigation benchmark without retraining, our multi-view variant raises success rate from 50% to 70%, and our LightGlue-style head raises SPL from 45.7 to 59.1.

Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev et al. · 0 citations
Review Jul 2026

Accuracy potential of visual localization exploiting high-end street-level imagery

Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicability and accuracy are often limited. Yet, the accuracy potential of visual localization has not been systematically investigated against survey-grade demands. This is mainly due to the lack of publicly available, large-scale outdoor datasets with ground-truth poses in the sub-centimeter range. In this work, we address both gaps. We introduce a scalable visual localization pipeline that employs precisely georeferenced, high-resolution street-level imagery directly as the scene representation. It combines prior-guided reference candidate selection with on-the-fly local Structure-from-Motion reconstruction and PnP-based pose estimation. We further present the FHNW Muttenz dataset, a real-world dataset covering a contiguous 10 km street network mapped in two mobile mapping campaigns approximately 1.5 years apart. It consists of high-resolution reference imagery and query sequences acquired by four different cameras across five representative scenes. All images are precisely co-registered, yielding 6-DoF ground-truth poses in the sub-centimeter range. Using this dataset, we evaluate the accuracy potential of visual localization. Our experiments demonstrate median pose accuracies in the range of 1-5 cm for translation and 0.05-0.1{\deg} for rotation, reaching as low as 1 cm and 0.03{\deg} under favorable conditions. These results show that visual localization can complement survey-grade GNSS positioning, paving the way for 3D geospatial data acquisition using consumer devices and fully automated georeferencing approaches. The dataset is publicly available at: https://fhnw-muttenz-vl-dataset.github.io/.

Jonas Meyer, S. Nebiker, P. Theiler et al. · 0 citations