Skip to content

SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation, achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.

Abstract

American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.

View source

Similar papers

Conference Aug 2026

A Greedy Skeleton Retrieval Framework for Vietnamese Text-to-Sign Generation

Sign Language Production (SLP) plays a crucial role in bridging the communication gap between the Deaf community and broader society, functioning alongside Sign Language Translation (SLT) and Recognition (SLR). In addition to the limited scale of available data, research on Vietnamese Sign Language (VSL) is further hin...

D. Thanh, Thang Cap · 0 citations
2026

Bridging Text-to-Sign Translation via Codebook-Oriented Pretraining

This work proposes a novel text-to-sign translation based on model pretraining, which enhances semantic alignment by inheriting codebook-oriented prior knowledge from masked self-supervised models.

Ninlawat Phuangchoke, C. Polprasert · 0 citations
Preprint Sep 2026

MoVT: Video-Augmented Motion Tokenizer for Text-to-Motion Generation

MoVT is introduced, a novel framework that effectively leverages the extensive range of human action videos to enhance text-to-motion generation and performs favorably against prior state-of-the-art methods across multiple key metrics.

Bei-Bei Jing, Tian-Le Guo, You-Jia Zhang et al. · 0 citations
Aug 2026

Variational Sign Language Translation

A novel framework based on conditional Variational autoencoder for SLT (VSLT) that facilitates direct and sufficient cross-modal alignment between sign language videos and spoken language text is proposed, and a shared Attention Residual Gaussian Distribution (ARGD) which considers the textual information as a residual...

Rui Zhao, Liang Zhang, Biao Fu et al. · 0 citations
Preprint Aug 2026

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional vi...

Eunjee Choi, J. Sung, Seongwhan Cho et al. · 0 citations
Preprint Aug 2026

SignRR: Retrieve and Refine Real Motion for Sign Language Production

Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand con...

Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.