Integrating Computer Vision and Large Language Models for Real-Time Blind Navigation
The emergence of artificial intelligence (AI) in assistive technology is a rapidly evolving field that offers a great opportunity for visually impaired people to gain greater freedom. Despite the high pace of development in AI, current solutions are still rather incoherent and expensive and lack contextual reasoning and multimodal interaction. The proposed research combines computer vision and speech processing with Internet of Things modules and large language model-based reasoning in a single hands-free device using a wearable, head-mounted design. The architecture is based on a Raspberry Pi-powered edge-computing platform and cloud-assisted AI services to support real-time perception of the environment, voice-based execution of tasks, and context-specific reactions. The experimental evaluation demonstrated a wake-word detection latency of 0.6 s and an overall response time of 3.8 s, 93% speech recognition in quiet conditions, and 88% object detection in typical lighting. A satisfaction score of 8.5/10 was obtained as a result of user testing with 10 participants. These findings indicate the practical feasibility of the proposed system for real-time assistive navigation, demonstrating low-cost implementation, contextual interaction capabilities, and the potential to address several limitations of existing wearable navigation systems.