Jul 2026· Engineering Research Express· Vol 8, pp. 145209· 0 citations· 28 references
Physics
TL;DR
A lightweight two-stream feature extractor, Recurrent Convolutional Recursive Transformer—Multi-Scale Convolution Augmented Transformer (RCRT-MCAT), for few-shot cross-domain HAR, which reaches a fine-tuned accuracy on the most challenging 76-way cross-environment gesture recognition task.
Abstract
Human activity recognition (HAR) using Wi-Fi channel state information (CSI) faces severe challenges in cross-domain generalization and data scarcity. Existing methods either rely on complex hardware deployment or suffer from insufficient spatiotemporal feature extraction, leading to poor performance under domain shifts. To address these issues, this paper proposes a lightweight two-stream feature extractor, Recurrent Convolutional Recursive Transformer—Multi-Scale Convolution Augmented Transformer (RCRT-MCAT), for few-shot cross-domain HAR. The model decouples CSI signals into a temporal stream and a channel stream to separately mine complementary spatiotemporal information. The MCAT branch employs multi-scale convolution and adaptive attention to capture fine-grained temporal patterns. The RCRT branch adopts recurrent convolution and recursive Transformer to efficiently model spatial dependencies across antennas and subcarriers. An evaluation framework is established on two public datasets, SignFi and Wiar, covering four experimental settings: in-domain recognition, cross-environment recognition, cross-environment cross-user recognition, and cross-dataset recognition. Experimental results demonstrate that, on the most challenging 76-way cross-environment gesture recognition task, the proposed model achieves an accuracy of 77.1% under the 1-shot setting, representing a 24.2 percentage point improvement over the FewSense model. When the number of samples is increased to 5-shot, the accuracy rises sharply to 91%, which is 28.2 percentage points higher than FewSense. In the cross-dataset recognition scenario, our model reaches a fine-tuned accuracy of 77.8% on user a2, 13.5% higher than FewSense. The average unfine-tuned accuracy across all users is 62.2%.
A cross-environment transfer learning framework for CSI-based HAR that integrates CSI preprocessing, adaptive amplitude-phase fusion via TinyGate, an R(2+1)D backbone, and two temporal modeling strategies, namely Bidirectional Long Short-Term Memory (Bi-LSTM) and Transformer is proposed.
The proposed UniSense-CSI is a unified multi-task framework that jointly learns dynamic gesture recognition, static posture classification, and fall detection and converts CSI signals into pseudo-RGB images, extracts spatio-temporal features using a customized ConvNeXt backbone, and leverages an improved PerceiverIO mo...
This paper proposes CGAC, a model that integrates convolutional bidirectional gated recurrent units with temporal attention, and shows that CGAC delivers the best performance on UT-HAR and remains competitive across different acquisition tools and CSI classification tasks.
Lili Cai· International journal of pat...· 0 citations
This paper proposes a novel hybrid Convolutional Neural Network-Gated Recurrent Unit (CNN-GRU) architecture for accurate channel state estimation in 2-user Non-Orthogonal Multiple Access (NOMA) systems operating under diverse wireless propagation environments. The proposed model integrates convolutional layers for effe...
S. Heshmat, Sara Khaled, Ahmed Ezzat et al.· IEEE Access· 0 citations
Conventional surveillance infrastructure relies almost exclusively on camera-based monitoring, which raises privacy concerns, requires adequate illumination and unobstructed line of sight, and depends on continuous human supervision. This paper presents C-Sense, a device-free, privacy-preserving framework for detecting...
Sourabh Nondi, Bornali Gogoi, Nelson R. Varte· International Journal of Res...· 0 citations
Wi-Fi Channel State Information (CSI) provides a privacy-preserving modality for human activity recognition (HAR), particularly in environments where activity classes vary in complexity and temporal scale. This work presents an integrated framework that combines multi-scale Fourier operator learning with manifold-aware...