Skip to content

MICA-Net: A Multimodal Cross-Attention Network for Human Action Recognition

Sep 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 61 references

TL;DR

A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.

Abstract

Automatic human action recognition (HAR) has become the most active research topic in recent years due to its broad applications, ranging from health monitoring and video surveillance to human-robot interaction. In this paper, we introduce a novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model. MICA-Net is composed of a lightweight 3D CNN, optimized for mobile devices, to extract visual features and a co-attention network that integrates a 2D CNN with a transformer to extract motion features. These extracted features are then continuously fused through dynamic gated Joint Cross-Attention Modules (JCAMs). These modules capture the intra- and inter- modal relationship while adaptively learning the contribution of each modality across different scenarios. We evaluated our recognition model on four publicly available multimodal datasets, MMAct, UESTC-MMEA-CL, UTD-MHAD, and MuWiGes. On MMAct, our model achieves an impressive F1-score of 89.24% with Cross-Subject and 96.57% with Cross-Session, outperforming current state-of-the-art methods. Similarly, on the UESTC-MMEA-CL, UTD-MHAD, and MuWiGes datasets, it achieves outstanding accuracy of 99.24%, 94.72%, and 98.98%, respectively. To demonstrate the practicality of the model, we design a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications. Real-time deployment indicates its potential for real implementation in the future. Our code is publicly available at https://anonymous.4open.science/r/MICA-Net-682E.

View source

Similar papers

Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.

Şafak Kılıç · 0 citations
2026

Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition

This study presents an efficient and interpretable deep learning framework for automatic handshake recognition from video sequences that employs a pretrained 3D Convolutional Neural Network to directly learn spatiotemporal features, enabling effective modelling of both motion dynamics and spatial relationships between...

S. Kumaravel, S. Veni · 0 citations
Conference Aug 2026

MSCALNet: a multiscale convolutional attention LSTM network for IMU-based human activity recognition

This work proposes MSCALNet, a Multi-Scale Convolutional Attention LSTM Network, a Multi-Scale Convolutional Attention LSTM Network that employs a multi-branch differential encoding strategy to fuse heterogeneous sensor information, and efficiently models multi-timescale dynamics through a dilated convolutional pyramid...

Zi-Bo Wang, Runyang Lyu, Bin Xiao · 0 citations
Open access Aug 2026

Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network

Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...

Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al. · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Open access Sep 2026

An open-set hybrid CNN-TCN network for human action recognition with unknown action rejection in collaborative robot workspaces

Introduction Open-set recognition capability is becoming necessary for collaborative robots because actual shop-floor dynamics do not always correspond to the fixed action classes used during training. Misclassification of unknown human motions may lead to inappropriate robot responses, whereas reliable rejection enabl...

Yasir Abdullah R, Ignisha Rajathi G, Barakkath Nisha U et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.