Skip to content
Conference

Keypoint-Based Isolated Sign Language Recognition via Multi-Stream MLP and Temporal Attention BiLSTM

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 712-717 · 0 citations · 29 references

Abstract

In this paper, we propose a keypoint-based isolated sign language recognition (ISLR) system addressing four challenges in large-vocabulary recognition: signer-dependent spatial variance, temporal dynamics, class imbalance, and multi-stream feature fusion. Using MediaPipe Holistic, we extract 1,662-dimensional skeletal keypoints covering pose, face, and hand dynamics. To achieve spatial invariance without additional trainable parameters, we introduce a parameter-free, shoulder-referenced normalization scheme that ensures approximate invariance under translation and scaling without distorting local geometric relationships. Class imbalance is mitigated through a leakage-free augmentation protocol of four targeted strategies, strictly confined to the training set, complemented by label smoothing and class-weighted loss. Built on these preprocessing foundations, three specialized MLP branches independently extract per-body-part features, which are concatenated and fed into a two-layer Bidirectional LSTM followed by a Temporal Attention mechanism that focuses on the most discriminative frames. Evaluated on a unified corpus of 309 sign classes aggregated from three public datasets, the proposed system achieves 97.20% accuracy and 89.58% macro F1 (mean over 3 seeds), outperforming several strong baselines while requiring fewer parameters and converging significantly faster. On the standard INCLUDE benchmark, the proposed model further achieves 97.76% accuracy under the 70/15/15 protocol and 95.76% under the 80/20 protocol, surpassing some prior published results.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.