Inclusive Offline Multimodal Retrieval-Augmented Generation System for Accessible PDF-Based Knowledge Assistance
Abstract
People with visual, speech and hearing impairment still face a major problem of receiving digital knowledge. In spite of the fact that the artificial intelligence enhances the information retrieval systems, the majority of the solutions are based on the cloud-based large language models and they do not offer an inclusive multimodal interaction. In this paper, an Offline Multimodal Retrieval-Augmented Generation (RAG) System is introduced that is intended to help differently-abled users to interact with PDF documents with the help of text, speech, and sign-language. The suggested system consists of the locally run large language model (LLaMA through Ollama), semantic retrieval based on FAISS, offline speech recognition, text-to-speech synthesis, and Sign animation rendering through gesture recognition. Our architecture is based on privacy, low latency and free deployment, unlike the traditional cloud-dependent AI assistants. Experimental evaluation demonstrates that the system effectively retrieves context-relevant responses while supporting voice interaction and visual gesture assistance. The proposed framework contributes toward inclusive AI-driven knowledge systems and demonstrates the feasibility of offline assistive intelligence platforms.