SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation, achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.
Abstract
American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.
Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...
D. Thanh, Thang Cap· International Conference on...· 0 citations
This work proposes a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models.
Ninlawat Phuangchoke, C. Polprasert· International Conference on...· 0 citations
MoVT is introduced, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation and performs favorably against prior state-of-the-art methods across multiple key metrics.
Bei-Bei Jing, Tian-Le Guo, You-Jia Zhang et al.· 0 citations
A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...
Rui Zhao, Liang Zhang, Biao Fu et al.· International Journal of Com...· 0 citations
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional vi...
Eunjee Choi, J. Sung, Seongwhan Cho et al.· 0 citations
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand con...
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.