Skip to content

Two-stream prototype network for Wi-Fi CSI-based cross-domain human activity recognition

Jul 2026 · Engineering Research Express · Vol 8, pp. 145209 · 0 citations · 28 references
Physics

TL;DR

A lightweight two-stream feature extractor, Recurrent Convolutional Recursive Transformer—Multi-Scale Convolution Augmented Transformer (RCRT-MCAT), for few-shot cross-domain HAR, which reaches a fine-tuned accuracy on the most challenging 76-way cross-environment gesture recognition task.

Abstract

Human activity recognition (HAR) using Wi-Fi channel state information (CSI) faces severe challenges in cross-domain generalization and data scarcity. Existing methods either rely on complex hardware deployment or suffer from insufficient spatiotemporal feature extraction, leading to poor performance under domain shifts. To address these issues, this paper proposes a lightweight two-stream feature extractor, Recurrent Convolutional Recursive Transformer—Multi-Scale Convolution Augmented Transformer (RCRT-MCAT), for few-shot cross-domain HAR. The model decouples CSI signals into a temporal stream and a channel stream to separately mine complementary spatiotemporal information. The MCAT branch employs multi-scale convolution and adaptive attention to capture fine-grained temporal patterns. The RCRT branch adopts recurrent convolution and recursive Transformer to efficiently model spatial dependencies across antennas and subcarriers. An evaluation framework is established on two public datasets, SignFi and Wiar, covering four experimental settings: in-domain recognition, cross-environment recognition, cross-environment cross-user recognition, and cross-dataset recognition. Experimental results demonstrate that, on the most challenging 76-way cross-environment gesture recognition task, the proposed model achieves an accuracy of 77.1% under the 1-shot setting, representing a 24.2 percentage point improvement over the FewSense model. When the number of samples is increased to 5-shot, the accuracy rises sharply to 91%, which is 28.2 percentage points higher than FewSense. In the cross-dataset recognition scenario, our model reaches a fine-tuned accuracy of 77.8% on user a2, 13.5% higher than FewSense. The average unfine-tuned accuracy across all users is 62.2%.

View source

Similar papers

Open access 2026

Low-Data Cross-Environment Transfer Learning for Wi-Fi CSI-Based Human Activity Recognition: An Inductive Bias Perspective

A cross-environment transfer learning framework for CSI-based HAR that integrates CSI preprocessing, adaptive amplitude-phase fusion via TinyGate, an R(2+1)D backbone, and two temporal modeling strategies, namely Bidirectional Long Short-Term Memory (Bi-LSTM) and Transformer is proposed.

Parma Hadi Rantelinggi, Mondher Bouazizi, Tomoaki Ohtsuki · 0 citations
Open access Aug 2026

Learning Shared Semantic Representations from WiFi CSI for Unified Multi-Task Human Activity Recognition

The proposed UniSense-CSI is a unified multi-task framework that jointly learns dynamic gesture recognition, static posture classification, and fall detection and converts CSI signals into pseudo-RGB images, extracts spatio-temporal features using a customized ConvNeXt backbone, and leverages an improved PerceiverIO mo...

Jing-Lun Mao, Qi-Yue Ma, Fengxia Han · 0 citations
Aug 2026

CGAC: A Convolutional Bidirectional GRU Network with Temporal Attention for WiFi CSI-Based Human Activity Recognition

This paper proposes CGAC, a model that integrates convolutional bidirectional gated recurrent units with temporal attention, and shows that CGAC delivers the best performance on UT-HAR and remains competitive across different acquisition tools and CSI classification tasks.

Lili Cai · 0 citations
Open access 2026

HY-CNN-GRU: A Universal Deep Learning Framework for Robust Channel Estimation Across Heterogeneous Wireless Environments in 6G NOMA Systems

This paper proposes a novel hybrid Convolutional Neural Network-Gated Recurrent Unit (CNN-GRU) architecture for accurate channel state estimation in 2-user Non-Orthogonal Multiple Access (NOMA) systems operating under diverse wireless propagation environments. The proposed model integrates convolutional layers for effe...

S. Heshmat, Sara Khaled, Ahmed Ezzat et al. · 0 citations
Open access Aug 2026

C-Sense: A Wi-Fi Channel State Information and LSTM-Based Deep Learning Framework for Privacy-Preserving Suspicious Human Activity Detection

Conventional surveillance infrastructure relies almost exclusively on camera-based monitoring, which raises privacy concerns, requires adequate illumination and unobstructed line of sight, and depends on continuous human supervision. This paper presents C-Sense, a device-free, privacy-preserving framework for detecting...

Sourabh Nondi, Bornali Gogoi, Nelson R. Varte · 0 citations
Open access Aug 2026

Adaptive Multi-Scale Fourier Neural Operator Learning with Manifold-Preserving Local Density Oversampling for Coarse-Grained Wi-Fi CSI-Based Human Activity Recognition

Wi-Fi Channel State Information (CSI) provides a privacy-preserving modality for human activity recognition (HAR), particularly in environments where activity classes vary in complexity and temporal scale. This work presents an integrated framework that combines multi-scale Fourier operator learning with manifold-aware...

Qiang Zhao, Yuchu Lin, Jia-Hui Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.