This work presents the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC), and leverages the decomposition of handshapes into five phonological features shared across both languages, to decode LSC handshapes from predicted features via a composite phonological distance metric.
Abstract
Sign language processing advances rapidly for high-resource languages such as American Sign Language (ASL), yet most of the world's sign languages lack the phonological annotations new methods require. We present the first zero-shot cross-lingual framework for handshape recognition, transferring from ASL to Catalan Sign Language (LSC). Our approach leverages the decomposition of handshapes into five phonological features -- selected fingers, flexion, spread, thumb position, and thumb contact -- shared across both languages, to decode LSC handshapes from predicted features via a composite phonological distance metric. We evaluate three architectures (MLP, SL-GCN, SHuBERT) trained on two ASL corpora (PopSign, Sem-Lex) against a 37-handshape, single-signer LSC benchmark. Zero-shot transfer proves viable once recording-format disparities are harmonized, reaching 80.0% phonological feature accuracy and 54.5% expected handshape accuracy. Phonological decomposition thus offers a bridge for extending sign language technologies to low-resource languages without any target-language video training labels.
We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (\'ITM). We use \'ITM SignWiki, a dataset derived from a bilingual Icelandic--\'ITM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the f...
F. Ingimundarson, Gudny Bjork Thorvaldsdottir, Mathias Müller et al.· 0 citations
A scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings, achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, attaining a BLEU-4 score of 10.03 and surpassing the previous best method by 0.67 points.
Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda et al.· Proceedings of the Thirty-Fi...· 0 citations
This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation that aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4.
Oline Ranum, Edward Fish, Simon Hadfield et al.· 0 citations
Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation systems. Advances in pose estimation, transformer architectures, and large-scale dataset collection have driven progress, yet challenges remain. Datasets are limited com...
Ozge Mercanoglu Sincan, A. Pelykh, Edward Fish et al.· 0 citations
This research introduces VSL400, a multi-view video dataset for isolated word-level recognition of Vietnamese Sign Language (VSL). The dataset contains 74,259 manually annotated video clips covering 400 glosses performed by 28 signers, including deaf student signers and additional trained VSL signers. Each signing in...
Trung Nguyen Quoc, Khoi Pham Dang, Viet Truong Duy et al.· Scientific Data· 1 citation
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.