Jul 2026· International Journal of Drug Delivery Technology· Vol 16· 0 citations· 11 references
TL;DR
An AI-based image generation system is introduced that utilizes a Stable Diffusion model fine-tuned with Low-Rank Adaptation (LoRA) for domain-specific image generation, demonstrating that SD+LoRA is an efficient and scalable domain-specific text-to-image generation system.
Abstract
Deep learning-based generative models have made a major leap forward in the world of image generation with the help
of Artificial Intelligence. One of the most notable of these developments is text-to-image synthesis, which can
automatically generate images based on natural language descriptions. In this work, an AI-based image generation system
is introduced that utilizes a Stable Diffusion model fine-tuned with Low-Rank Adaptation (LoRA) for domain-specific
image generation. The main idea of the proposed system is to combine the text encoding of CLIP, the latent compression
of Variational Autoencoder (VAE), and the denoising ability of diffusion to create images that are both semantically
relevant and visually coherent based on text prompts. The proposed approach was tested on a Pokemon image-caption
dataset for fine-tuning the pre-trained Stable Diffusion model and its effectiveness evaluated. The study shows that the
diffusion-based architectures outperform the traditional GAN based methods in terms of image quality, training stability,
semantic alignment, and output diversity. The main advantage of LoRA fine-tuning was the substantial decrease in
computational load, which involved updating just a small fraction of trainable parameters without compromising the
model's performance. Experimental results indicated that successful images of Pokemon could be generated, and that the
images were consistent with the text attributes such as color, type, and appearance. The results demonstrate that SD+LoRA
is an efficient and scalable domain-specific text-to-image generation system. The research underscores the rising
significance of diffusion-based generative AI in digital content creation, imaginative design, entertainment, and cleverness
in visual generation systems
The paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models are discussed.
Image caption generation is a significant area of study in artificial intelligence and computer vision, focusing on training systems to generate accurate textual descriptions of images. This paper presents an advanced framework for image captioning using a dataset sourced from Kaggle. The process begins with pre-proces...
Bhargavi Polepalli, Praveen Kumar Sekharamantry, K. S. Rao· Pertanika journal of science...· 0 citations
This study investigates the automatic classification of real and AI-generated flower images using fine-tuned transfer learning models and shows that Swin Transformer-Tiny achieved the best overall performance, reaching an F1-score of 88.64% and outperforming the other architectures.
Mehtap Ülker· NATURENGS MTU Journal of Eng...· 0 citations
The proposed Attention-Based Deep Learning Pipeline of AI-Created Image Recognition incorporates three integrated branches, including low-level statistical feature extraction, high-level semantic representation learning, and attention-based feature refinement mechanism, which support the robustness and generalization a...
Nadia Ali· Al-Noor Journal of Engineeri...· 0 citations
Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The cen...
Yushun Tang, Weiming Chen, Siyi Liu et al.· IEEE transactions on multime...· 0 citations
It is shown that diffusion models also inherently use an attention mechanism very similar to that of transformers, and similarities in basic functional principle of auto-encoders and attention-based models allows for interchange of designs based on practical requirements.
F. Haddadi, L. Monfared, Ebrahim Rezaii et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.