Skip to content
Conference

A Greedy Skeleton Retrieval Framework for Vietnamese Text-to-Sign Generation

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 682-687 · 0 citations · 23 references

Abstract

Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hindered by the absence of standardized gloss annotations. Glossing serves as the syntactic representation of sign language; it is an essential intermediary because spoken and signed languages possess fundamentally different grammatical structures, cognitive logic, and word orders. To address these challenges, we propose a lightweight and effective framework for low-resource text-to-sign generation. Specifically, the proposed pipeline first employs a pretrained Vietnamese Transformer encoder-decoder model to translate spoken sentences into intermediate gloss sequences to bridge this linguistic gap. Subsequently, leveraging a data-driven dictionary that maps glosses to skeletal representations, we introduce an efficient greedy-based retrieval strategy to extract and match the corresponding skeletal segments. Finally, these retrieved pose sequences are seamlessly synthesized into continuous sign language videos. Experimental results demonstrate that the proposed framework is effective in a low-resource setting, achieving strong performance in text-to-gloss translation and improved coverage in gloss-to-skeleton mapping compared to conventional tokenization-based approaches. In addition, the interpolation strategy significantly enhances motion smoothness during rendering. These findings suggest that the proposed method provides a practical baseline for text-to-sign generation in under-resourced sign languages.

View source

Similar papers

2026

Bridging Text-to-Sign Translation via Codebook-Oriented Pretraining

This work proposes a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models.

Ninlawat Phuangchoke, C. Polprasert · 0 citations
Aug 2026

Variational Sign Language Translation

A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...

Rui Zhao, Liang Zhang, Biao Fu et al. · 0 citations
Conference Open access Sep 2026

A Gloss-driven Indian Sign Language Production System Using Learned Pose Representations

A scalable and modular SLP framework based on Sign-Pose-VQ-VAE model, designed for low-resource settings, achieves state-of-the-art performance among keypoint-based methods on the PHOENIX14T benchmark, attaining a BLEU-4 score of 10.03 and surpassing the previous best method by 0.67 points.

Suvajit Patra, Arkadip Maitra, Swami Punyeshwarananda et al. · 0 citations
#computer vision Preprint Sep 2026

SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale

Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unifi...

Zhaoyi An, Si-Han Tan, Youngbae Hwang et al. · 0 citations
Preprint Aug 2026

SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs

Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be...

Shi-Wei Gan, Xiao Liu, Ya-Feng Yin et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.