Pyramid Target Perception Network with Efficient Context Modeling and Multi-Scale Cross-Attention for Infrared Small Target Detection
Abstract
Infrared small target detection (IRSTD) is a challenging task in intelligent infrared sensing and electronic imaging systems, because dim targets often occupy only a few pixels and are easily disturbed by clutter, noise, and low-contrast background structures. A practical detector should preserve pixel-level target cues while suppressing target-like false responses. This paper proposes a Pyramid Target Perception Network (PTPN) for single-frame pixel-level IRSTD. The network integrates three complementary components: an Efficient Context Modeling (ECM) encoder employing 7 × 7 depthwise separable convolution for lightweight contextual feature extraction, a multi-scale target cross-attention (MTCA) module for hierarchical feature interaction, and a small-target feature pyramid network (STFPN) for target-preserving multi-scale aggregation. In addition, a physics-constrained loss (PCL) is introduced during training to regularize predictions according to infrared imaging characteristics, including point spread consistency, target-region relative intensity consistency, and signal-to-noise-ratio-aware separability. Experiments on IRSTD-1k, NUAA-SIRST, and NUDT-SIRST demonstrate that PTPN achieves IoU scores of 71.87%, 79.56%, and 86.47%, respectively, with 4.55M parameters, 4.96G FLOPs at an input resolution of 256 × 256, and an inference speed of 45.0 FPS. Although PTPN achieves competitive overall performance, it does not attain the highest IoU on NUDT-SIRST, indicating that pixel-level target-region estimation under complex scenes remains an area for further improvement. Overall, PTPN provides an effective balance between target localization, false-alarm suppression, and computational efficiency, supporting its potential application in AI-driven infrared image processing and intelligent electronic sensing systems.