Skip to content
Preprint

Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

Abstract

We present B2TFPose, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images. Using a single frozen DINOv3 vision transformer as its only pretrained component within the pose estimation pipeline, B2TFPose extracts dense patch-level features that generalize across the synthetic-to-real domain gap without any task-specific fine-tuning, revisiting the classical local feature matching paradigm through the lens of large-scale self-supervised foundation models. Three contributions advance the training-free state of the art. A geodesic non-maximum suppression strategy retrieves a viewpoint-diverse template set for coarse-to-fine correspondence matching. Render-guided Re-Correspondence (RRC) synthesizes object-specific views at the estimated pose and re-establishes dense 2D-3D correspondences to sharpen the initial estimate without additional learned parameters. A multi-mask hypothesis selection strategy jointly scores competing segmentation candidates to resolve segmentation ambiguity. On the seven core datasets of the BOP Benchmark, B2TFPose achieves 40.7 mean AR without refinement and 56.4 with refinement, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

View source

Similar papers

Preprint Aug 2026

Foundational feature fusion for conditional flow matching in 6D pose estimation

This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.

Amir Hamza, Davide Boscaini, Fabio Poiesi · 0 citations
Preprint Sep 2026

LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation

Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...

Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al. · 0 citations
Conference Aug 2026

SNOA-Pose: Symmetry-Aware Refinement for Sim-to-Real Category-Level 6D Object Pose Estimation

Category-level 6D object pose estimation recovers the rotation, translation, and scale of unseen object instances within specific categories from RGB-D observations. Voting-based methods such as CPPF++ can be trained without real pose annotations. However, the gap between clean CAD-based training samples and noisy, inc...

Jia-Jun Xiong, Qian Min, Dong Li · 0 citations
Preprint Sep 2026

GRC-Pose: Generation-Reconstruction Correspondence for Prior-Free 6D Object Pose Tracking

Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, posed reference images, or pose annotations. Geometric foundation models provide complementary object-centric and scene-centric cues, yet SAM3D CAD is indexed by an arbitrar...

Shi-Yang Liu, Wei-Quan Lin, Lu-Ping Xiao et al. · 0 citations
Preprint Sep 2026

DIFTA-3D: Depth-Consistent Instance-Level Feature Transfer and Adaptation of DINOv3 for 3D Detection

RGB-D 3D instance detectors benefit from visual semantics, but the task-specific Faster R-CNN/ResNet branch used by IIFNet3D couples feature extraction to a separately trained 2D detector and its image-domain labels. Replacing that branch with a frozen vision foundation model removes this task-specific dependency, but...

Lin-Man Wang, Zi-Fei Zhang, Chun-Ran Zheng et al. · 0 citations
Preprint Oct 2026

Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encode...

Xin-Hao Xiang, Wei-Yang Li, Zhi-Jie Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.