Skip to content
Conference

Visible and infrared image registration based on deep geometric similarity evaluation

Sep 2026 · Ninth Global Intelligent Industry Conference (GIIC 2026) · 0 citations

Abstract

To address the difficulty of accurately aligning target regions in visible and infrared images caused by differences in imaging mechanisms and inconsistent radiometric discrepancies, a multi-stage registration method based on a deep geometric similarity evaluation network is proposed. The method constructs a structure-preserving modality translation generator and a geometric similarity evaluator, and synthesizes pseudo cross-modal samples through geometric perturbation to disentangle appearance discrepancies from structural misalignment, enabling the geometric similarity evaluator to learn response representations that are sensitive to geometric offsets while being robust to modality differences. In the registration stage, a two-stage framework combining global affine coarse registration and local deformable fine registration is adopted. Moreover, the geometric similarity evaluation is innovatively used as a unified supervision signal throughout both the global and local optimization processes, effectively improving the consistency of optimization objectives across different scales. Experiments were conducted on two visible-infrared cross-modal datasets, MRSR and LLVIP. The results show that the proposed method achieves the best performance on both datasets. On the MRSR remote sensing dataset, the Reprojection Error (RE) is reduced to 4.027, the Corner Root Mean Square Error (RMSEcor) reaches 0.701, and the Displacement Field Root Mean Square Error (RMSEdisp) is 0.0086. On the LLVIP low-light target dataset, the above three metrics reach 3.075, 0.437, and 0.0087, respectively. Compared with current state-of-the-art joint global and local registration methods, the proposed method reduces the RE and RMSEcor on the MRSR dataset by approximately 32.0% and 36.1%, respectively; on the LLVIP dataset, the RMSEcor is further reduced by 52.7%. The proposed method demonstrates strong robustness and generalization capability in both visible-infrared remote sensing observation and low-light target perception scenarios.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.