Lightweight Dual Radar Point Cloud Learning for Cross Scene Human Behavior Recognition
Abstract
Millimeter-wave radar enables privacy-preserving human behavior recognition, but point-cloud observations remain sparse, viewpoint-dependent and sensitive to scene changes. This paper presents a lightweight dual-radar point-cloud learning method for cross-scene behavior recognition. The two radar streams are first registered into a unified coordinate system and temporal ly fused to reduce instantaneous sparsity. A compact frame-level point encoder extracts spatial descriptors from four-dimensional radar points, and a shallow temporal Transformer models motion evolution across frames. Experiments on 4,848 paired dual-radar samples covering eight behaviors show that the proposed method achieves stable cross-scene recognition, with an Accuracy of 0. 8586 at K=5 and a Macro-F1 of 0.8538 at K=11. Robustness analysis further indicates that random perturbations cause limited degradation, whereas structured spatial masking is the dominant failure case. The lightweight baseline contains 1.44M parameters and reaches 440.0 FPS, showing a favorable trade-off between accuracy and deployment efficiency.