MSP-Edge: Multi-Stage Pruning With Neuron Reconstruction for Privacy-Aware Edge Vision-Based Crowd Sensing in IoT-Enabled Rail Transit Systems
Abstract
Internet-of-Things (IoT)-enabled surveillance systems increasingly utilize edge intelligence to support real-time analytics while reducing network dependency and limiting the transmission of raw visual data, thereby reducing potential data exposure. In rail transit environments, crowd monitoring requires low-latency on-device visual sensing that remains robust under dense crowds, severe occlusion, motion blur, and complex illumination conditions. However, deploying deep neural detectors on resource-constrained edge platforms remains challenging due to limited computational resources and the difficulty of translating pruning-induced sparsity into practical acceleration. This paper presents MSP-Edge, a deployment-aware compression framework that combines multi-stage unstructured sparsification during model optimization with deployment-oriented neuron reconstruction for efficient and privacy-aware edge vision-based crowd sensing. MSP-Edge integrates dynamic suppression, TL1/DTL ${_{q}}$ regularization, mask-guided importance tracking, one-shot pruning, and post-pruning stabilization to progressively remove redundant parameters while preserving detection performance. Instead of directly deploying the resulting irregular sparse network, a neuron reconstruction mechanism reorganizes computational dimensions rendered inactive by sparsification and compacts the surviving computation into dense tensor layouts, improving memory locality and embedded inference efficiency while preserving approximate functional consistency with the pruned model. Unlike conventional channel pruning, reconstruction is performed only after sparsification has converged, introduces no additional pruning decisions, and serves specifically as a deployment-oriented graph transformation. Experimental results show that MSP-Edge reduces model weights by 58.2% and floating-point operations by 61.7% relative to YOLOv8n while maintaining comparable detection accuracy on the RPEE-Heads dataset. A controlled ablation further shows that the stabilized sparse model and reconstructed MSP-Edge model retain the same 58.2% sparsity and identical 82.7 mAP@0.5, while reconstruction reduces median latency from 42 ms to 31 ms, demonstrating that the additional deployment acceleration originates primarily from reconstruction rather than further parameter removal. Deployment on an NVIDIA Jetson Xavier NX using TensorRT FP16 achieves a 32.6% reduction in median latency and a 23.6% reduction in p99 latency relative to the uncompressed YOLOv8n baseline. Cross-dataset evaluations and multi-seed experiments further demonstrate consistent detection performance across the evaluated crowd datasets and stable optimization behavior across repeated training runs. These results demonstrate that MSP-Edge provides an effective deployment pipeline that transforms sparsified neural networks into compact dense models for low-latency and privacy-aware crowd sensing under the evaluated edge deployment conditions. Here, “privacy-aware” specifically refers to architectural data minimization through local edge inference and reduced raw visual-data transmission rather than formal privacy-preserving guarantees.