SeRoNet: Semantic-Aware Channel Routing for UAV Visible-Infrared Object Detection
Abstract
Visible-infrared uncrewed aerial vehicle (UAV) object detection is essential for robust all-day aerial perception, as visible imagery provides rich appearance cues while infrared imagery offers stable thermal responses under adverse illumination. However, existing multimodal detectors mainly rely on feature-level interaction, which does not necessarily ensure semantic consistency across modalities. In complex UAV scenes, small objects, cluttered backgrounds, sensor misalignment, and modality-specific distortions may cause corresponding targets to exhibit divergent feature distributions, while the reliability of each modality also varies across channels and imaging conditions. To address these issues, we propose SeRoNet, a semantic-aware and selective fusion framework for visible-infrared UAV object detection. In particular, the semantic prior extraction module (SPEM) constructs offline semantic prototypes using a frozen vision-language model (VLM), and the semantic constraint module (SCM) regularizes detector features toward modality-invariant semantic representations during training without introducing additional inference cost. Meanwhile, the cross-modal channel routing (CMCR) module performs channelwise selective interaction to transfer reliable complementary information while suppressing redundant or unreliable responses. Extensive experiments on the DroneVehicle and vehicle detection in aerial imagery (VEDAI) benchmarks demonstrate that SeRoNet achieves 84.6% and 84.7% mean average precision (mAP), respectively, outperforming state-of-the-art methods while maintaining competitive computational efficiency. Code is available at: https://github.com/skytoyo/SeRoNet