Skip to content
Preprint

XDG: Accelerated Visual Disambiguation

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

XDG is introduced, an efficient visual disambiguation model designed for scalable SfM that remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup.

Abstract

Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically distinct surfaces can produce incorrect image matches and degrade reconstruction quality. Previous work mitigates this issue with geometry-aware foundation-model features, but places a heavy transformer classifier on top of the backbone, making large-scale disambiguation expensive. We introduce XDG, an efficient visual disambiguation model designed for scalable SfM. Our key observation is that a 3D foundation model already performs the cross-view geometric reasoning necessary for visual disambiguation, so doppelganger classification should adapt the backbone representation directly rather than relearn pair reasoning in a separate heavy decoder. XDG fine-tunes Depth Anything 3 with lightweight LoRA adapters and repurposes its camera tokens as compact pair-level classification tokens. A compact MLP head predicts whether a candidate image pair observes the same 3D surface. Extensive experiments show that XDG provides a favorable accuracy-efficiency tradeoff: it remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup. On individual LaMAR scenes containing thousands of images, XDG saves more than 10 hours of visual disambiguation processing. Code is available at https://github.com/xtcpete/xdg.

View source

Similar papers

Preprint Aug 2026

ID-VTG: Image-Disambiguated Video Temporal Grounding

The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.

Minghang Zheng, Jing Wei, Hong-Yi Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding

Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current Vision Language Models (VLMs) lack. VLMs built on CLIP-based ViT backbones trade fine-grained detail for high-level semantics, and we show this loss propagates downstream. To recover it, we fuse DINOv3 and Cle...

Harsha Patnala, Debopriyo Banerjee, A. Munot et al. · 0 citations
Preprint Aug 2026

RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars

Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coarse to fully specify visual identity, so they fail to separate visually similar distractors. Few-shot counters sidestep this with visual exemplars, but require manual an...

A. d'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh · 0 citations
#computer vision Preprint Aug 2026

TRACE: Transparent Retrieval for Abstract Concept Evaluation

Recent work reports that vision--language models (VLMs) struggle to establish and maintain stable reference in repeated reference games. Rather than ask which VLM does best, we ask a more basic question: do you need a large pretrained VLM for this at all? On grounding a single director utterance to one of twelve tangra...

Joe M. Bingham · 0 citations
Preprint Aug 2026

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

G2D is proposed, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image and transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning.

Zehua Hao, Fang Liu, Qinliang Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-...

Guan-Xu Yu, Yu-Hang Yao · 5 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.