Smart Sight: A Comprehensive Deep Learning Framework for Real-Time Assistive Navigation and Object Recognition for Visually Impaired Individuals
Visual impairment is a universal health burden, and one of the most pressing concerns in the world, as it is estimated that there are 2.2 billion afflicted persons in the world, which is highly constraining in their free movements, their interactions with the surrounding, as well as on their social interrelation. White canes and guide dogs, which are some of the classical assistive solutions, may only offer the simplest contextual awareness and lack the semantic awareness of the scene to crawl safely and reliably through the complex real-world world. In this paper, I will introduce Smart Sight, a custom real-time assistive navigation architecture that tightly integrates four existing state-of-theart deep learning models, such as YOLOv8nano to identify objects, Deep-SORT to track multiple objects, MiDaS v3.0 to estimate object and face depth, and a priority-based Text-to-Speech (TTS) engine to offer intelligent feedback, and none The proposed system achieves an average of 94.3% of the Mean Average Precision (mAP) on the MS COCO data benchmark and 90.9% on a non-test dataset of three classes in the context of road-anomaly and can continue to provide end-to-end inference of 22 frames per second with an aggregate latency of 145 milliseconds on consumer mobile hardware.