Signsight: An Edge-AI Based Multimodal Assistive Framework for OCR, Isolated Sign Recognition, and Scene-to-Speech Feedback
Abstract
EdgeAssist is a multimodal Edge-AI Assistive Framework with integrated sign language recognition, OCRbased text extraction, scene-aware object detection and speech interaction in a single local application to support accessibility. The framework uses MediaPipe hand landmark extraction, a Random Forest hand sign recognizer for isolated hand sign classification, and Tesseract OCR for printed-text recognition and YOLOv8 Nano for real-time scene understanding. The outputs recognized are transformed to speech by a light-weight text to speech engine, thus allowing for an easy, intuitive humancomputer interaction for visually and hearing-impaired users. The proposed system does not rely on the cloud; therefore, it is not latency-sensitive and it is more private. Experimental results show that the system works well in real time, can respond to the user in time, and provides multimodal support well in normal indoor environments. The structure offers a feasible and scalable base for future intelligent assistive technologies.