This work proposes XPos3R, a generalizable pose regression method that eliminates preoperative preparation, and introduces an asymmetric encoder-decoder architecture that improves cross-modal feature alignment while maintaining computational efficiency.
Abstract
Intraoperative 2D/3D registration, which aligns live X-ray images with preoperative volumes, is essential for image-guided interventions. Previous regression-based methods suffer from limited generalization, thus requiring time-consuming patient-specific retraining. Inspired by recent geometry foundation models such as DUSt3R, we propose XPos3R, a generalizable pose regression method that eliminates preoperative preparation. Unlike existing geometry models designed for homogeneous inputs, XPos3R extends this paradigm to multi-modal inputs, namely 2D X-rays and 3D volumes. Specifically, we introduce an asymmetric encoder-decoder architecture that improves cross-modal feature alignment while maintaining computational efficiency. To scale training under limited medical data, we adopt an anatomy-specific data curation strategy and construct million-scale synthetic datasets. Evaluated on real-world benchmarks, a single pretrained XPos3R surpasses patient-specific methods in both accuracy and robustness. With test-time optimization completed in seconds, it further reduces the 3D error to<4 mm and the reprojection error to<1 mm. The strong generalization, accuracy, and efficiency of XPos3R highlight its clinical potential, while its asymmetric framework may inspire broader cross-modal vision geometry tasks.
Aligning intraoperative biplanar digital subtraction angiography (DSA) to pre-procedural computed tomography angiography (CTA) requires rapid and accurate 3D-to-2D registration. Optimization-based methods are sensitive to initialization and may require hundreds of iterations, whereas learning-based approaches commonly...
R. L. M. van Herten, Robert Graf, Paula Feldman et al.· 0 citations
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue motion. Conventional multi-view 3D reconstruction methods rely on stable correspondences and approximate rigidity, which a...
Zhi-Hao Xing, Ying-Yu Wang, Liang Zhao et al.· 0 citations
Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks...
Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures often exhibit substantially different image intensities across modalities. In this work, we present a comprehensive benchmark of intra-patient 3D multimodal deformable regi...
Matteo Barbieri, G. La Barbera, J. De la Plata et al.· 0 citations
A simple and efficient self-supervised pre-training framework for 3D medical images based on a two-fold patch-wise perturbation strategy, requiring substantially less memory, computation, and training time than the state-of-the-art pre-training pipelines.
Tirthajit Baruah, Kabir Jamadar, Punit Rathore· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.