Sep 2026· IEEE Internet of Things Journal· Vol 13, pp. 40324-40337· 1 citation· 43 references
Computer Science
Abstract
The proliferation of multimicrophone IoT devices—such as smartphones, smart speakers, and wearables—has enabled diverse acoustic sensing applications. However, scarcely labeled data and poor cross-task transferability continue to limit traditional supervised methods. To address this, we propose acoustic-URL, an unsupervised representation learning framework using cross-domain fusion (cdf) and cross-channel contrastive learning (CCCL) to extract transferable features from unlabeled multichannel audio. Our approach integrates time-domain waveforms and spatial-domain channel impulse responses (CIRs) to capture both temporal and spatial information. By aligning features from different microphones that capture the same acoustic event, the model learns spatially consistent representations. Evaluations on four tasks—digit recognition, writer identification, gesture recognition, and footstep-based identification—under varying supervision and task-domain shifts show that acoustic-URL consistently outperforms supervised baselines, especially in low-label settings, and demonstrates strong cross-task generalization.
Results indicate that multi-scale Inception encoding and auxiliary supervision are the most important components, while the PatchX interaction module provides no consistent additional gain under the current setting.
Tian-Chang Xie, Hai-Ling Wang, Wei-Guang Wang et al.· IEEE Photonics Journal· 0 citations
Fiber-optic distributed acoustic sensing (DAS) has emerged as a critical Internet-of-Things (IoT) sensing technology with broad industrial applications. However, the two-dimensional spatial-temporal morphology of DAS signals presents analytical challenges for conventional methods. In contrast, this complex signal struc...
J. Duan, Jia-Geng Chen, Zuyuan He· Science China Information Sc...· 0 citations
This work frames this as Domain-Incremental Learning over acoustic environments and presents the first continual learning benchmark for multi-channel replay speech detection, evaluating a state-of-the-art beamformer-based detector over all 24 environment orderings of the ReMASC corpus with five seeds.
Voice Activity Detection (VAD) is a fundamental component that supports a wide range of audio/speech processing applications. Numerous studies have addressed VAD using a single microphone or a compact microphone array, yet their performance remains limited in distant scenarios. In this paper, we develop a novel data-dr...
De Hu, Shao-Jie Li, Qing-Ying Zhao et al.· IEEE Transactions on Audio,...· 0 citations
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
Results establish MADS (Multi-view Acoustic Descriptor Set) not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.
Utsab Ghosh, Roshni Chakraborty· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.