2026· IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing· Vol 19, pp. 28192-28206· 0 citations· 58 references
TL;DR
This work designs a Feature Representation Refinement module that stabilizes RGB and IR features through scale-specific channel remapping and normalized nonlinear refinement, yielding more reliable representations for subsequent cross-modal interaction and develops a cross-modal feature interaction mechanism.
Abstract
RGB–infrared (RGB-IR) vehicle detection in uncrewed aerial vehicle (UAV) imagery is essential for applications, such as traffic monitoring and object tracking. However, existing methods often suffer from heterogeneous feature responses across modalities, degraded RGB feature representations under adverse illumination, and insufficient capture of fine-grained structural and edge details, which collectively impede accurate cross-modal modeling and localization. To alleviate these issues, we propose an Illumination-aware Semantic-Guided Mamba (ISGM) network. First, we design a Feature Representation Refinement module that stabilizes RGB and IR features through scale-specific channel remapping and normalized nonlinear refinement, yielding more reliable representations for subsequent cross-modal interaction. Furthermore, to enhance robustness to illumination variations and better preserve detailed structural cues and boundary information, we develop a cross-modal feature interaction mechanism comprising the Illumination-Aware Fusion Modulation (IAFM) module and the Detail-enhanced Semantic-Guided Mamba (DSGM) module. Specifically, the IAFM module estimates illumination-aware modality reliability weight maps, thereby improving robustness under challenging illumination conditions. These weight maps guide the DSGM module to integrate high-level semantic information and low-level detail cues into multiscale guidance features. These features are subsequently used to modulate the scanning parameters for adaptive RGB-IR feature interaction. This design improves the modeling of target-region features while minimizing interference from complex backgrounds. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate that ISGM outperforms state-of-the-art methods in detection performance while achieving a favorable accuracy-efficiency balance among comparable methods.
This paper proposes a Difference-guided Detail Enhancement Fusion Network (DDEF-Net) for UAV-based RGB–thermal (RGB-T) object detection, which enables effective complementary exploitation of visible and infrared information in complex scenarios. A Difference-guided Kolmogorov–Arnold Network (KAN) Calibration Fusion mod...
Yu-Jie Li, Zheng-Sheng Chen, D. Ma et al.· Remote Sensing· 0 citations
In modern smart agriculture, the uncrewed aerial vehicles (UAVs) have emerged as indispensable low-altitude remote sensing platforms, particularly for precision irrigation and crop protection. While they enable the efficient acquisition of high-resolution spatial information and significantly enhance agricultural produ...
Yi-Chen Liu, Wen-Xuan Zhang, Xuan-Rui Chen et al.· IEEE Transactions on Geoscie...· 0 citations
A Boundary-Guided Dual-Perspective Cross-Modal Fusion Network (BDPNet) is proposed to explicitly preserve shallow geometric structures and decouple deep semantic fusion into macroscopic and microscopic perspectives.
Hu Lin, Zhi-Wei Fu, Xiu-Mei Chen et al.· Remote Sensing· 0 citations
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object de...
Cun-Zheng Fan, Dawei Yan, Guan-Lin Wang et al.· 0 citations
Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially varying modality reliability, while the small visual extent...
Qi-Fan Zhang, Zi-Ran Zhou, Rui-Jie Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.