Skip to content
Conference

Enhanced YOLOv11 with global geometric perception for 6D object pose estimation

Jul 2026 · International Conference on Image Processing and Intelligent Control · Vol 14262, pp. 142620K - 142620K-7 · 0 citations · 11 references
Engineering

Abstract

Precise 6D object pose estimation from RGB images remains a formidable challenge due to complex backgrounds and severe occlusions. To address these issues, our study presents an enhanced YOLOv11 framework specifically designed for global geometric perception and high-fidelity single-stage 6D pose regression. The core of our architecture is the C3k2SW module, which innovatively synergizes local convolutional features with global long-range dependencies through windowbased self-attention, significantly enhancing the network's geometric perception of spatial topologies. Furthermore, to optimize multi-scale feature interaction, an adaptive ConcatA module and a Bi-directional Feature Pyramid Attention Network (BFPAN) are proposed to suppress background noise while preserving fine-grained geometric details across different scales. Experimental results on the LineMod benchmark demonstrate that our method achieves an optimal tradeoff between inference efficiency and accuracy, reaching an average ADD(-S) accuracy of 76.50% and 84.92% on the 5cm 5° metric, respectively. These results validate that the integration of global geometric awareness consistently outperforms the vanilla YOLOv11 and other classical baselines in complex scenarios.

View source