This work presents FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoders supervised on object-scene overlap and reducing supervision requirements and memory overhead.
Abstract
Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial feature fusion strategies to resolve pose ambiguities. We present FunFlow6D, a novel flow matching-based formulation that leverages features from geometric and appearance foundation models for pose estimation, eliminating the need for task-specific encoder training. We also introduce a cross attention-based fusion mechanism that dynamically combines geometric and appearance features to provide richer conditioning for the flow matching module. Experiments on four datasets from the BOP benchmark show that FunFlow6D outperforms the previous state of the art while reducing supervision requirements and memory overhead. Extensive ablations validate the contribution of each proposed component. Project website: https://tev-fbk.github.io/FunFlow6D/.
B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.
Robust 6D object pose estimation is essential for enabling intelligent systems to interact with the physical world. While existing methods have achieved notable progress, they remain constrained by two major limitations: poor scalability when new objects are introduced and severe performance degradation caused by catas...
Long Tian, Yang Liu, Jun-Lin Fang et al.· IEEE Transactions on Image P...· 0 citations
RGB-D 3D instance detectors benefit from visual semantics, but the task-specific Faster R-CNN/ResNet branch used by IIFNet3D couples feature extraction to a separately trained 2D detector and its image-domain labels. Replacing that branch with a frozen vision foundation model removes this task-specific dependency, but...
Lin-Man Wang, Zi-Fei Zhang, Chun-Ran Zheng et al.· 0 citations
Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encode...
Xin-Hao Xiang, Wei-Yang Li, Zhi-Jie Zheng et al.· 0 citations
Category-level 6D object pose estimation recovers the rotation, translation, and scale of unseen object instances within specific categories from RGB-D observations. Voting-based methods such as CPPF++ can be trained without real pose annotations. However, the gap between clean CAD-based training samples and noisy, inc...
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object...
Hong-Li Xu, Zhao-Wei Lu, Jun-Wen Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.