Skip to content

A Multi-Modal Deep Learning Framework for Assistive Vision: Software-Based Design and Real-Time Validation

· 0 citations · 14 references

TL;DR

The development of an intelligent visual aid incorporating object detection, face recognition, distance estimation, distance estimation, and text-to-speech audio feedback is outlined.

View source

Similar papers

Conference Aug 2026

Smart Sight: A Comprehensive Deep Learning Framework for Real-Time Assistive Navigation and Object Recognition for Visually Impaired Individuals

Visual impairment is a universal health burden, and one of the most pressing concerns in the world, as it is estimated that there are 2.2 billion afflicted persons in the world, which is highly constraining in their free movements, their interactions with the surrounding, as well as on their social interrelation. White canes and guide dogs, which are some of the classical assistive solutions, may only offer the simplest contextual awareness and lack the semantic awareness of the scene to crawl safely and reliably through the complex real-world world. In this paper, I will introduce Smart Sight, a custom real-time assistive navigation architecture that tightly integrates four existing state-of-theart deep learning models, such as YOLOv8nano to identify objects, Deep-SORT to track multiple objects, MiDaS v3.0 to estimate object and face depth, and a priority-based Text-to-Speech (TTS) engine to offer intelligent feedback, and none The proposed system achieves an average of 94.3% of the Mean Average Precision (mAP) on the MS COCO data benchmark and 90.9% on a non-test dataset of three classes in the context of road-anomaly and can continue to provide end-to-end inference of 22 frames per second with an aggregate latency of 145 milliseconds on consumer mobile hardware.

S. M. Raj, V. Hemanth, N. Nikhil et al. · 0 citations

AI-Based Object Detection and Smart Navigation System for the Visually Impaired

An AI-powered assistive system designed to enhance the mobility, safety, and independence of visually impaired individuals, creating a smart, voice-guided companion that empowers visually impaired users to navigate their surroundings with confidence and independence.

G. Sireesha, K. Prasanthi, Kanchumarthi Nirmala et al. · 0 citations
Open access Sep 2026

Integrating Computer Vision and Large Language Models for Real-Time Blind Navigation

The emergence of artificial intelligence (AI) in assistive technology is a rapidly evolving field that offers a great opportunity for visually impaired people to gain greater freedom. Despite the high pace of development in AI, current solutions are still rather incoherent and expensive and lack contextual reasoning and multimodal interaction. The proposed research combines computer vision and speech processing with Internet of Things modules and large language model-based reasoning in a single hands-free device using a wearable, head-mounted design. The architecture is based on a Raspberry Pi-powered edge-computing platform and cloud-assisted AI services to support real-time perception of the environment, voice-based execution of tasks, and context-specific reactions. The experimental evaluation demonstrated a wake-word detection latency of 0.6 s and an overall response time of 3.8 s, 93% speech recognition in quiet conditions, and 88% object detection in typical lighting. A satisfaction score of 8.5/10 was obtained as a result of user testing with 10 participants. These findings indicate the practical feasibility of the proposed system for real-time assistive navigation, demonstrating low-cost implementation, contextual interaction capabilities, and the potential to address several limitations of existing wearable navigation systems.

Ashish Sharma, Dhiraj Rajput, Sushant Kumar et al. · 0 citations
Open access 2026

Vision-Enabled Virtual Robotic Head with Multi-Modal Communication Capability.

AI-enabled intelligent virtual humanoids are critical to ensure natural and socially aware human-machine interactions in areas such as education, customer services, and companionship. In this paper, we develop a vision-enabled virtual robotic head which is entirely running inside a web browser and includes modules like Speech-to-Text, Large Language Models, Text-to-Speech, and vision processing using FastAPI and WebSocket architecture. The proposed approach generates a dynamic robot-like face capable of realistic talking and facial expressions along with a vision module for face detection and visitor recognition and personalized interaction based on their memory. The proposed architecture ensures low-latency (less than one second) despite the active use of vision processing capabilities. User evaluation on 50 people resulted in 92% face recognition accuracy and high satisfaction scores.

Md. Ashiqussalehin, M. Khushi, Zannatul Ferdushie et al. · 0 citations
Open access 2026

MultiModal Deep Learning Framework for Missing Children Detection using Vision-Language Mode

As report of missing children continue to rise around the world there is an urgent requirement of some smart and scalable systems which can help in timely and accurate identification. In this paper, we propose a multimodal deep learning system that combines Vision-Language Transformers like CLIP, Face Re-identification networks, and Graph Neural Networks (GNNs) in order to help in locating and tracking missing children. The proposed system is unlike conventional facial recognition systems, which can only act on visual stimulus whereas the proposed system handles a multitude of different stimulus e.g. facial image, descriptive text, metadata location, time greater and context of the stimulus e.g. description of clothing. The vision language module enables visual inputs to be correlated to the textual description so that the model is able to interpret and correlate the narrative indicators within the public reports. A face re-identification module with high tolerances to visual variability because of age, lighting, occlusion, etc. and a GNN to learn spatial and temporal links between sightings to make movement pattern predictions. Initial tests on publicly available data show a better matching across the unimodal baselines and a substantial improvement in recall and contextual relevance in the metrics. Uniting the visual, linguistic, and structural data, the model provides more accurate findings in real-life situations and demonstrates a high possibility of combining it with police technology. There should be more work on increasing child-targeted multimodal data sets and evaluating in the real world testing.

K. U. Maheswari, P. Dhanalakshmi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.