Overcoming Label Scarcity in Distributed Acoustic Sensing via Masked Spectrogram Modeling
Abstract
Distributed acoustic sensing (DAS) enables continuous monitoring of large-scale infrastructure, but the deployment of robust intelligent systems is severely hindered by the scarcity of expertly annotated data, cross-environmental domain shifts, and spatial biases caused by irregular physical fiber geometries. To address these bottlenecks, we propose masked spectrogram modeling (MSM), a physics-informed self-supervised learning (SSL) framework tailored specifically for DAS. Rather than treating DAS data as a rigid spatiotemporal matrix, MSM employs a channel-independent design to achieve topology invariance and reduce sensitivity to real-world deployment distortions. The framework utilizes a highly efficient Vision Transformer pretrained on a pretext task of reconstructing heavily masked, joint time–frequency patches from massive volumes of unlabeled short-time Fourier transform (STFT) spectrograms. We validate MSM on two complex, real-world industrial datasets—a metro construction site and a gas pipeline—and extensive ablation studies rigorously justify the domain-specific architectural adaptations. The results demonstrate that MSM learns acoustic representations that support both same-deployment adaptation and cross-deployment transfer: on the Shenzhen metro intrusion dataset (SMID) metro benchmark, MSM reaches 96.17% accuracy with 500 labeled samples per class and 95.40% with 100 labeled samples per class, while, on the Gas pipeline monitoring dataset (GPMD) pipeline transfer benchmark, it reaches 94.20% with 500 labeled samples per class and remains above 91% with 100 labeled samples per class. Furthermore, the optimized shallow architecture requires only 1.84M parameters, 3.36M floating-point operations (FLOPs) per inference, and 40.71-MB peak VRAM, providing a highly scalable and label-efficient solution for industrial DAS applications.