Skip to content

Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition

2026 · International Journal of Latest Technology in Engineering, Management & Applied Science · Vol 15, pp. 722-732 · 0 citations

TL;DR

This study presents an efficient and interpretable deep learning framework for automatic handshake recognition from video sequences that employs a pretrained 3D Convolutional Neural Network to directly learn spatiotemporal features, enabling effective modelling of both motion dynamics and spatial relationships between interacting individuals.

Abstract

Human Activity Recognition (HAR) has gained significant attention in computer vision due to its wide range of applications in surveillance, social behaviour analysis, and human–computer interaction. Among various human-to-human interactions, handshake recognition is particularly important as it represents social intention and cooperative behaviour. This study presents an efficient and interpretable deep learning framework for automatic handshake recognition from video sequences. The proposed approach employs a pretrained 3D Convolutional Neural Network (3D CNN) to directly learn spatiotemporal features, enabling effective modelling of both motion dynamics and spatial relationships between interacting individuals. The experiments are conducted using two dataset namely UT-Interaction Human Interaction Dataset and SBU Kinect Interaction dataset, focusing exclusively on the handshake interaction as the target class. The dataset provides accurate ground-truth annotations, including temporal intervals and bounding boxes, which support precise localization and reliable recognition of handshake actions. Each dataset is split into 80% for training, 10% for validation, and 10% for testing to ensure robust performance evaluation. The experimental results demonstrated that the proposed 3D CNN-based framework achieved a handshake recognition accuracy of 98.92% on the UT-Interaction dataset, representing performance improvements of 9.72%, 6.32%, and 7.12% compared to CNN, BiLSTM, and RNN models, respectively.

View source

Similar papers

Open access Aug 2026

Recognition and Modeling of User Interaction Behaviors in Virtual Reality Based on 3D Convolutional Neural Networks

Accurate recognition of continuous user interactions in virtual reality (VR) is essential for intelligent perception systems and provides methodological support for multimodal information processing in electromagnetic sensing and nextgeneration human–machine interaction. However, fragmented action sequences, weak long-...

H.-L. Wu · 0 citations
Sep 2026

MICA-Net: A Multimodal Cross-Attention Network for Human Action Recognition

A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.

Trung-Hieu Le, Thai-Khanh Nguyen, T. Tran et al. · 0 citations
Conference Aug 2026

A Hybrid Deep Learning Approach for Emotion Recognition Using Gait Patterns, Body Gestures, and Facial Expressions

Human emotion recognition from visual cues has gained significant attention due to its applications in human-computer interaction, surveillance, and healthcare. However, relying on a single modality often limits the robustness of emotion understanding in real-world scenarios. To address this, we propose a novel hybrid...

R. J. R. Kumar, G. Brindha, S. Gayathri et al. · 0 citations
Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.

Şafak Kılıç · 0 citations
Open access Aug 2026

Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network

Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...

Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al. · 0 citations
Aug 2026

Enhancing facial expression recognition with lightweight attention-based CNN

A lightweight FER model that combines a truncated MobileNetV2 backbone with a patch-based local feature extraction module and a channel-attention refinement module and a channel-attention refinement module, followed by a compact classifier is proposed, outperforming many existing lightweight models.

Anh Viet Vu, Khanh Gia Pham, Nam Quy Tran et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.