An AI-powered assistive system designed to enhance the mobility, safety, and independence of visually impaired individuals, creating a smart, voice-guided companion that empowers visually impaired users to navigate their surroundings with confidence and independence.
The development of an intelligent visual aid incorporating object detection, face recognition, distance estimation, distance estimation, and text-to-speech audio feedback is outlined.
Visual impairment is a universal health burden, and one of the most pressing concerns in the world, as it is estimated that there are 2.2 billion afflicted persons in the world, which is highly constraining in their free movements, their interactions with the surrounding, as well as on their social interrelation. White canes and guide dogs, which are some of the classical assistive solutions, may only offer the simplest contextual awareness and lack the semantic awareness of the scene to crawl safely and reliably through the complex real-world world. In this paper, I will introduce Smart Sight, a custom real-time assistive navigation architecture that tightly integrates four existing state-of-theart deep learning models, such as YOLOv8nano to identify objects, Deep-SORT to track multiple objects, MiDaS v3.0 to estimate object and face depth, and a priority-based Text-to-Speech (TTS) engine to offer intelligent feedback, and none The proposed system achieves an average of 94.3% of the Mean Average Precision (mAP) on the MS COCO data benchmark and 90.9% on a non-test dataset of three classes in the context of road-anomaly and can continue to provide end-to-end inference of 22 frames per second with an aggregate latency of 145 milliseconds on consumer mobile hardware.
S. M. Raj, V. Hemanth, N. Nikhil et al.· International Conference on...· 0 citations
Visually impaired people (VIP) need assistive technology for smooth and confident navigation on pavements along the roads. VIPs find it hard to follow the speech generated through a text-to-audio converter due to the presence of advertisement text. Automating the process of audio conversion from the signboards by integrating the capabilities of emerging AI models in a fraction of a second will be useful for their confident navigation. In this context, SignSpeak, an AI-based assistive utility, is designed to convert signboard text to speech using visual inputs, which may be captured through smart sensors or canes. The proposed system follows a three-step (detection-extraction-conversion) pipeline. Initially, a fine-tuned lightweight YOLOv11s model is used to generate a bounded box for the relevant text region, which is input to Gemini model, a multimodal large language model for text extraction. Lastly, the extracted text is converted into audio using the Google Text-to-Speech tool. We fine-tuned the YOLOv11s model on 200 signboard images and tested its performance on an additional curated dataset of 42 real-world signboard images. Experimental findings indicate that the proposed integrated pipeline achieved superior performance compared to the solitary usage of the Gemini model for text extraction followed by its audio conversion.
Sharanjit Kaur, Manju Bhardwaj, Kriti Misra et al.· ITEGAM- Journal of Engineeri...· 0 citations
Within the framework of assistive technologies for visually impaired users, this paper introduces JU-EYE robot, a smart simulation-based navigation system that is meant to be used on the University of Jordan campus. The proposed system combines a depth-enabled camera (RGB-D) with a YOLO detection model to categorize obstacles while also using depth data to estimate their distances and a scene sensing layer linked to an AI-assistant to generate context-aware decision making. The system uses graph algorithms to plan paths, which help users to find the best and safest routes around the campus independently. By providing real-time audio feedback through an AI-based assistant supported with bilingual voice interaction, the system provided impactful navigation and robust scene awareness.
Farah Elyan, Lina Issa, Kareman Marouf et al.· IEEE Jordan Conference on Ap...· 0 citations
A hybrid navigation framework is proposed to improve mobility assistance for visually impaired users by incorporating GPS-based turn-by-turn guidance with real-time vision-based obstacle detection. Current navigation systems offer directional guidance but are not aware of dynamic environmental obstacles, making them less suitable for real-world pedestrian navigation. The limitation is addressed by a multimodal solution, which uses voice-assisted GPS navigation and computer vision. The navigation component allows users to input locations via voice inputs and receive turn-by-turn instructions in real-time using GPS. The vision component uses real-time object detection with YOLOv8, while depth estimation is achieved with Depth Pro. A semantic risk model is also introduced to evaluate obstacles based on the type of objects, distance, and location in the path. The environment is segmented into navigational regions, and a least-risk direction is dynamically determined. The system also ensures safety by pausing navigation guidance when obstacles with high risk are detected and giving immediate audio feedback. Experimental results demonstrate effective obstacle detection with over 90% accuracy and real-time system response below 1.5 seconds, validating the system’s suitability for practical deployment. The framework can help improve safe navigation and situational awareness for the visually impaired.
Tummala Adithya Reddy, Thushara Thampi, Malavika Sreejith et al.· 2026 7th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.