SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.
Abstract
Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.
The results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.
Chauhan Pareshbhai Mansangbhai, D. Vaghela, Mahesh Goyani et al.· International Journal of Ele...· 0 citations
In this paper, we propose a keypoint-based isolated sign language recognition (ISLR) system addressing four challenges in large-vocabulary recognition: signer-dependent spatial variance, temporal dynamics, class imbalance, and multi-stream feature fusion. Using MediaPipe Holistic, we extract 1,662-dimensional skeletal...
Thi Kim Ngan Tran, Dung Ha Nguyen, Thanh Binh Nguyen· International Conference on...· 0 citations
SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.
Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden· 0 citations
Continuous Sign Language Recognition (CSLR) is a challenging sequence-to-sequence task that requires simultaneous modeling of frame-level motion patterns, gloss-level transitions, and sentence-level dependencies. While traditional CSLR methods mainly emphasize frame-level feature extraction, they often insufficiently c...
Ngoc Quoc Tran, Thi Diem Tran· International Conference on...· 0 citations
Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D...
A. Pelykh, Edward Fish, Ozge Mercanoglu Sincan et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.