Aug 2026· International Conference on Advanced Sensing and Intelligent Systems· Vol 14309, pp. 143090W - 143090W-8· 0 citations· 12 references
Engineering
TL;DR
This paper proposes a novel Transformer-based neural network architecture specifically designed for radar signal processing that integrates multi-head self-attention mechanisms with temporal convolutional networks to effectively model both local patterns and global dependencies in radar data.
Abstract
Radar-based target classification and intent recognition play crucial roles in autonomous driving, surveillance, and aerospace applications. Traditional methods relying on handcrafted features and conventional deep learning architectures often fail to capture long-range dependencies in radar sequences. This paper proposes a novel Transformer-based neural network architecture specifically designed for radar signal processing. Our approach integrates multi-head self-attention mechanisms with temporal convolutional networks to effectively model both local patterns and global dependencies in radar data. We design a hierarchical feature extraction module that processes range-Doppler maps and micro-Doppler signatures for robust target classification. Furthermore, we introduce an intent recognition module that leverages trajectory prediction and behavioral pattern analysis. Extensive experiments on simulated and real radar datasets demonstrate that our method achieves 94.7% classification accuracy and 89.3% intent recognition accuracy, outperforming state-of-the-art methods including CNNbased and LSTM-based approaches. The proposed architecture also shows superior computational efficiency with 35% fewer parameters compared to baseline models.
With the increasing complexity of the electromagnetic environment, traditional radar target recognition methods face severe challenges. High-resolution range profile (HRRP) and Radar Cross Section (RCS), as two important radar features, each has its own advantages in target recognition but also exhibits limitations. To enhance radar target recognition performance in complex scenarios such as low signal-to-noise ratio (SNR), this paper proposes a recognition method based on heterogeneous multi-modal feature fusion. The proposed method constructs a three-channel parallel encoding network, which utilizes Convolutional Long Short-Term Memory (ConvLSTM), One-Dimensional Convolutional Gated Recurrent Unit (Conv1D-GRU), and Gated Recurrent Unit (GRU) to extract deep discriminative features from raw HRRP sequences, RCS sequences, and HRRP statistical features, respectively. Furthermore, it innovatively designs a dual-path cooperative fusion mechanism, achieving explicit inter-modal correlation modeling through a cross-attention module and dynamically learning the importance of each modality through an adaptive weight fusion layer, thereby realizing deep complementarity and enhancement of multi-modal information. Experimental results demonstrate that under various signal-to-noise ratios and polarization conditions, the proposed method achieves a maximum average recognition accuracy of over 99% for 6 ship targets. Compared with the traditional three-channel fixed-weight fusion method, the recognition accuracy of the proposed method increases from 90.01% to 99.42% under co-polarization, and from 81.14% to 98.15% under cross-polarization, fully validating the effectiveness and superiority of the dual-path fusion mechanism.
Sifan Su, Wei Yang, Shiwen Lei et al.· Remote Sensing· 0 citations
Human Activity Recognition (HAR) supports intelligent healthcare, surveillance, assisted living, and human–machine interaction. Vision-based methods are often limited by privacy concerns, illumination changes, and occlusion. This study proposes a hybrid spatial–temporal deep learning framework for FMCW radar-based HAR using micro-Doppler spectrograms. Four architectures are compared: 3D CNN–LSTM, 3D Bi-LSTM–CNN, CNN–Dilated Convolution–LSTM, and a Hybrid Ensemble CNN-LSTM with a Decision Tree classifier. Radar processing includes beat-frequency extraction, Range FFT, Doppler FFT, clutter suppression, and spectrogram generation. Convolutional layers extract spatial features, while LSTM and Bi-LSTM networks model temporal dependencies; dilated convolution expands the receptive field efficiently. Experimental results show that the hybrid models outperform conventional CNN and standalone LSTM approaches in accuracy, robustness, and generalisation. The hybrid ensemble achieves the best performance by combining spatial–temporal learning with ensemble optimisation while remaining effective in noisy environments and preserving user privacy.
Daffa Ahmadhan Khusumah, F. Suratman, Hesty Susanti· ARMADA : Jurnal Penelitian M...· 0 citations
Synthetic Aperture Radar (SAR) images always provide high-resolution data in all weather and lighting circum-stances. However, speckle noise, clutter backgrounds and scale variation remain as significant challenges for accurate target detection on SAR images. This paper proposes a hybrid deep learning method, which combines Convolutional Neural Networks (CNN), STDNet model and CFAR-based detection, with the objective to enhance performance of multi-scale object detection. The outputs from both the CNN and STDNet branches are fused using an Intersection over Union (IoU)-based fusion strategy helps to achieve better detection accuracy. The experimental results show that the proposed hybrid model reaches better precision, recall, and F1-score than separate classifiers. This proposed approach is a cost-effective and practical strategy for tracking different real-world SAR targets.
Nagamani Divedari, Kusma Kumari Cheepurupalli, Srinivasa Rao Chanamallu et al.· Journal of Intelligent Decis...· 0 citations
The synthetic aperture radar (SAR) object detection is crucial for military reconnaissance and environmental monitoring. However, the existing methods often struggle to maintain high accuracy in complex scenarios due to severe speckle noise, large variations in target scale, and similar feature interference. To address these challenges, this letter proposes SV-DAPNet, a robust detection framework based on the faster R-CNN architecture. We introduce three key innovations: 1) spatial variance modulation (SVM), which quantifies spatial variance to suppress speckle noise and enhance target features adaptively, 2) dynamic atrous spatial pyramid pooling (DASPP), which dynamically fuses multiscale features to handle large-scale variations, and 3) dynamic prototype contrastive learning (DPCL), which optimizes feature distribution to improve intraclass compactness and interclass discrimination. Extensive experiments on the SAR-AIRcraft-1.0 aircraft dataset and HRSID ship dataset demonstrate that SV-DAPNet achieves 89.6% mAP@0.5 and 49.7% mAP@0.5:0.95 on SAR-AIRcraft-1.0, outperforming the state-of-the-art baselines, and delivers consistent performance gains on cross-dataset validation with reasonable parameter consumption.
Ziheng Xia, Jinhua Wei, Wenjun Huo et al.· IEEE Geoscience and Remote S...· 0 citations