Jul 2026· 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS)· pp. 1682-1688· 0 citations· 15 references
Abstract
People with physical and motor disabilities face significant barriers when interacting with image creation applications that rely on keyboard input, touchscreens, or complex graphical interfaces. Although modern AI image generation models produce high-fidelity images from textual descriptions, they remain dependent on typed input, which is inaccessible to users with limited hand mobility. To address this gap, this paper proposes a voice-to-image assistive system that enables users to create images exclusively through voice commands. The system integrates three tightly coupled modules: (i) a Speech-to-Text (STT) engine based on OpenAI Whisper, which transcribes spoken commands into raw text; (ii) an NLP-based prompt refinement module that applies tokenization, grammatical correction, and semantic augmentation to produce diffusion-friendly image prompts; and (iii) a Stable Diffusion image synthesis backend that generates high-resolution images from the refined prompts using GPU-accelerated latent diffusion. The system achieves STT accuracy of up to 100% for simple commands and 96% for long descriptive inputs, with prompt-to-image semantic alignment reaching up to 97%. Average image generation time is 6–12 seconds on GPU hardware. This work demonstrates the practical viability of multimodal AI pipelines for assistive applications and outlines directions for future development, including multilingual support, mobile deployment, and integration with existing assistive technology ecosystems.
Blind and Low Vision Individuals (BLV) encounter significant difficulties in comprehending complex visual environments, while current assistive technologies typically lack speech-interactive reasoning capabilities. To address this gap, this study fine-tunes the Qwen2.5-Omni framework using Llamafactory based on our enh...
Li-Ting Chen, Ping-Ting Lin, Wei-Wei Li et al.· International Conference on...· 0 citations
Human-computer interaction has progressed significantly in recent years, yet existing digital assistants continue to suffer from limitations in emotional engagement, personalisation, and expressive communication. Most current AI systems operate primarily through text or voice, lacking a visual presence that allows user...
A. S. Sundhar, J. A. Jeba, V. Ramkumar et al.· FMDB Transactions on Sustain...· 0 citations
Gesture-based interaction represents a natural and intuitive paradigm for human-computer interaction (HCI), yet current systems struggle to bridge the semantic gap between low-level gesture recognition and high-level interface adaptation. This paper proposes GAUID (Gesture-Aware UI Design), a novel framework that lever...
Le Li· International Conference on...· 0 citations
People with visual, speech and hearing impairment still face a major problem of receiving digital knowledge. In spite of the fact that the artificial intelligence enhances the information retrieval systems, the majority of the solutions are based on the cloud-based large language models and they do not offer an inclusi...
B. V., Balavigneshwaran. SN., M. B. et al.· International journal of re...· 0 citations
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...
: The synthesis of photorealistic human faces from descriptive inputs has become feasible with the advancement of generative artificial intelligence. This paper presents C72: AI-Driven Realistic Human Face Creation, a framework that converts eyewitness free-text descriptions or guided attribute inputs into structured J...
H. M., M. K· Proceedings of the 1st Inter...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.