CFIMNet: Cross-Modal Frequency-Guided Iterative Multiscale Feature Fusion Network for Multimodal Remote Sensing Object Detection
Abstract
Multimodal object detection can improve detection accuracy and robustness by fusing complementary information from different modalities. However, in remote sensing scenarios, complex environments often weaken data reliability and cross-modal complementarity, making fixed-pattern fusion strategies insufficient for robust and accurate detection. To address these challenges, this article proposes CFIMNet, a cross-modal frequency-guided iterative multiscale feature fusion (IMFF) network that improves feature complementarity and multiscale representation for RGB–infrared remote sensing object detection. Specifically, a cross-modal frequency enhancement (CFE) module is designed to model channel–spatial interactions between RGB and infrared features and enhance structural information and fine-grained details through adaptive frequency decomposition and reconstruction. In addition, an IMFF module is introduced to progressively fuse cross-modal features across multiple scales within a parameter-shared iterative framework, thereby improving semantic consistency while preserving modality-specific characteristics. Extensive experiments on three benchmark remote sensing datasets demonstrate that CFIMNet consistently outperforms the state-of-the-art methods in detection accuracy and robustness. Code will be available at https://github.com/MapleMoonY/CFIMNet