SignAssistant: An AI-Enabled Real-Time Gesture-to-Speech System for Indian Sign Language Translation
Abstract
Communication barriers between deaf and hard-ofhearing (DHH) individuals and the hearing population remain a pressing global challenge, compounded in India by a severe shortage of certified Indian Sign Language (ISL) interpreters. This paper presents SignAssistant, a real-time, AI-enabled gesture-tospeech system that translates ISL alphabets, word-level gestures, and Hindi (Devanagari) character signs into multilingual text and speech. The system combines MediaPipe-based hand-landmark extraction with a hybrid CNN–Random Forest classifier for static gestures and an attention-augmented LSTM for dynamic signs, followed by multilingual translation into English, Hindi, and Marathi and neural text-to-speech synthesis. Unlike prior systems, which are typically restricted to a single gesture category, a single output language, or offline processing, SignAssistant unifies alphabet-, word-, and Hindi-letter recognition within one real-time pipeline and further enables sign-based querying of large language models. Evaluated on curated ISL and Hindi sign datasets, the hybrid model achieved 96-99% accuracy for alphabets, 90-93% for words, and 92-94% for Hindi letters, while sustaining approximately 15 frames per second and sub-200 ms latency on commodity CPU hardware. These results indicate that SignAssistant offers a scalable, multilingual, and interactive assistive technology that meaningfully advances accessibility for the DHH community.