An Infrared and Visible Image Fusion-Based Method for Air-to-Ground Target Tracking
Abstract
To address the challenges of low illumination, occlusion, complex backgrounds, platform motion, and dense distributions of small-scale targets in UAV-based air-to-ground missions, this paper investigates a target tracking method based on infrared and visible image fusion. Based on the complementarity of dual-modal imaging, MoME-Track is constructed for continuous single-target locking, while DCAF-Net/MAMC-Track is developed for multi-target detection and tracking. For single-target tracking, dynamic collaboration between appearance information and motion priors is achieved through cross-modal appearance experts, an extended Kalman filtering-based motion expert, and a mixture-of-experts decision mechanism. For multi-target tracking, a dual-branch cross-domain fusion detection network is designed to extract infrared and visible features. Frequency-spatial collaborative enhancement and multi-head cross-attention are introduced to improve small target detection capability. In the tracking stage, multimodal appearance measurement, depth-adaptive Kalman filtering, and low-confidence detection reuse are incorporated to enhance trajectory continuity and identity consistency. Experiments conducted on the self-constructed MSOT-UAV and MMOT-UAV datasets, as well as public datasets, demonstrate that the proposed method achieves a favorable balance among accuracy, robustness, and real-time performance.