Skip to content
Conference

HOPE-Based Temporal Modeling for Continuous Sign Language Recognition

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 736-741 · 0 citations · 25 references

Abstract

Continuous Sign Language Recognition (CSLR) is a challenging sequence-to-sequence task that requires simultaneous modeling of frame-level motion patterns, gloss-level transitions, and sentence-level dependencies. While traditional CSLR methods mainly emphasize frame-level feature extraction, they often insufficiently capture dynamic temporal relationships across video frames. Recent CSLR systems have shown that effective temporal modeling is crucial, with Transformer-based temporal encoders demonstrating strong performance in capturing both local and global dependencies. In this paper, we revisit temporal modeling for CSLR by replacing the Transformer-based encoder with HOPE, a recent architecture introduced under the Nested Learning paradigm. HOPE incorporates memory and update mechanisms that enable richer contextual information flow, making it well suited for modeling both local sign dynamics and long-range sentence-level dependencies. Combined with an efficient visual backbone, our proposed HOPE-CSLR forms an end-to-end recognition pipeline and achieves promising results, with WERs of 19.1% and 19.6% on the RWTH and RWTHT datasets, respectively. These results highlight the potential of memory-based temporal modeling for sign language recognition.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.