Skip to content

SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

Sep 2026 · 0 citations · 68 references
Computer Science

TL;DR

SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.

Abstract

Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.

View source

Similar papers

Open access Aug 2026

Enhancing Continuous Sign Language Recognition through MGPT-based Segmentation and Structured Position-Aware Decoding

The results indicate that introducing motion-consistent segmentation and structured decision fusion seems to be a good way for updating the CSLR systems beyond simply endwise paradigms.

Chauhan Pareshbhai Mansangbhai, D. Vaghela, Mahesh Goyani et al. · 0 citations
Conference Aug 2026

Keypoint-Based Isolated Sign Language Recognition via Multi-Stream MLP and Temporal Attention BiLSTM

In this paper, we propose a keypoint-based isolated sign language recognition (ISLR) system addressing four challenges in large-vocabulary recognition: signer-dependent spatial variance, temporal dynamics, class imbalance, and multi-stream feature fusion. Using MediaPipe Holistic, we extract 1,662-dimensional skeletal...

Thi Kim Ngan Tran, Dung Ha Nguyen, Thanh Binh Nguyen · 0 citations
Preprint Sep 2026

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning, surpassing methods explicitly trained on BSL and outperforming prior skeleton-based methods.

Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden · 0 citations
Conference Aug 2026

HOPE-Based Temporal Modeling for Continuous Sign Language Recognition

Continuous Sign Language Recognition (CSLR) is a challenging sequence-to-sequence task that requires simultaneous modeling of frame-level motion patterns, gloss-level transitions, and sentence-level dependencies. While traditional CSLR methods mainly emphasize frame-level feature extraction, they often insufficiently c...

Ngoc Quoc Tran, Thi Diem Tran · 0 citations
Preprint Sep 2026

SignRefine: Adapting Foundational Video Models for Sign Language Generation

Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D...

A. Pelykh, Edward Fish, Ozge Mercanoglu Sincan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.