Skip to content
Conference

An AI-Driven Voice-to-Image Generation Model for Assistive Applications

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 1682-1688 · 0 citations · 15 references

Abstract

People with physical and motor disabilities face significant barriers when interacting with image creation applications that rely on keyboard input, touchscreens, or complex graphical interfaces. Although modern AI image generation models produce high-fidelity images from textual descriptions, they remain dependent on typed input, which is inaccessible to users with limited hand mobility. To address this gap, this paper proposes a voice-to-image assistive system that enables users to create images exclusively through voice commands. The system integrates three tightly coupled modules: (i) a Speech-to-Text (STT) engine based on OpenAI Whisper, which transcribes spoken commands into raw text; (ii) an NLP-based prompt refinement module that applies tokenization, grammatical correction, and semantic augmentation to produce diffusion-friendly image prompts; and (iii) a Stable Diffusion image synthesis backend that generates high-resolution images from the refined prompts using GPU-accelerated latent diffusion. The system achieves STT accuracy of up to 100% for simple commands and 96% for long descriptive inputs, with prompt-to-image semantic alignment reaching up to 97%. Average image generation time is 6–12 seconds on GPU hardware. This work demonstrates the practical viability of multimodal AI pipelines for assistive applications and outlines directions for future development, including multilingual support, mobile deployment, and integration with existing assistive technology ecosystems.

View source

Similar papers

Conference Aug 2026

VizAdapt: novel dataset development with a voice-interactive visual question answering system for blind and low vision individuals

Blind and Low Vision Individuals (BLV) encounter significant difficulties in comprehending complex visual environments, while current assistive technologies typically lack speech-interactive reasoning capabilities. To address this gap, this study fine-tunes the Qwen2.5-Omni framework using Llamafactory based on our enh...

Li-Ting Chen, Ping-Ting Lin, Wei-Wei Li et al. · 0 citations
Open access Aug 2026

A Multimodal Transformer-Based Digital Companion for Emotion-Aware Human–Computer Interaction Using Real-Time 3D Avatars

Human-computer interaction has progressed significantly in recent years, yet existing digital assistants continue to suffer from limitations in emotional engagement, personalisation, and expressive communication. Most current AI systems operate primarily through text or voice, lacking a visual presence that allows user...

A. S. Sundhar, J. A. Jeba, V. Ramkumar et al. · 0 citations
Conference Aug 2026

Gesture-aware adaptive interface design using vision-language models for enhanced human–computer interaction

Gesture-based interaction represents a natural and intuitive paradigm for human-computer interaction (HCI), yet current systems struggle to bridge the semantic gap between low-level gesture recognition and high-level interface adaptation. This paper proposes GAUID (Gesture-Aware UI Design), a novel framework that lever...

Le Li · 0 citations
Open access 2026

Inclusive Offline Multimodal Retrieval-Augmented Generation System for Accessible PDF-Based Knowledge Assistance

People with visual, speech and hearing impairment still face a major problem of receiving digital knowledge. In spite of the fact that the artificial intelligence enhances the information retrieval systems, the majority of the solutions are based on the cloud-based large language models and they do not offer an inclusi...

B. V., Balavigneshwaran. SN., M. B. et al. · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
Conference Open access 2025

AI-Driven Realistic Human Face Creation from Natural Language Prompts

: The synthesis of photorealistic human faces from descriptive inputs has become feasible with the advancement of generative artificial intelligence. This paper presents C72: AI-Driven Realistic Human Face Creation, a framework that converts eyewitness free-text descriptions or guided attribute inputs into structured J...

H. M., M. K · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.