XGMR: Detector-Free RGB-Thermal Image Registration With Geometric Inductive Bias
Abstract
RGB-thermal image registration is essential for multi-sensor perception, yet accurate alignment remains difficult because RGB and thermal images differ substantially in appearance, gradient structure, and noise characteristics. Even fine-tuning a strong detector-free matcher like LoFTR on cross-modal pairs produces no valid matches at the 0.2 confidence threshold used at evaluation, and ground-truth registration labels are scarce in cross-modal settings. To address this problem, we propose XGMR, a detector-free framework for RGB-thermal image registration that incorporates geometric inductive bias while enabling fully self-supervised learning. Specifically, during training XGMR adds a homography-guided bias to cross-attention and complements it with an equivariance-based self-supervised loss for geometric consistency, without any labeled data. At inference the model runs global cross-attention without an explicit bias mask. A Modality Bridging Adapter narrows the spectral gap, a Dual-Softmax layer enforces bidirectional matching, a learned Self-Calibrating Head replaces RANSAC for homography refinement, and a Quality-Aware Fusion module gates the output by local registration quality. We evaluate XGMR on three public RGB-thermal datasets (NII-CU MAPD, LLVIP, and MSRS) covering aerial, nighttime, and daytime road scenes. XGMR achieves the lowest mean corner error among the compared methods on two datasets but trails XoFTR on LLVIP, where XoFTR’s KAIST pretraining provides a same-domain advantage. The full model uses 4.80M parameters and runs at 18.9 FPS on a single GPU. Combining geometry-guided self-supervision with a learned post-processing head registers cross-modal images without labeled data.