Skip to content
Open access

AI-Driven Image Synthesis from Textual Descriptions Using Stable Diffusion

Jul 2026 · International Journal of Drug Delivery Technology · 0 citations · 11 references

Abstract

Deep learning-based generative models have made a major leap forward in the world of image generation with the help of Artificial Intelligence. One of the most notable of these developments is text-to-image synthesis, which can automatically generate images based on natural language descriptions. In this work, an AI-based image generation system is introduced that utilizes a Stable Diffusion model fine-tuned with Low-Rank Adaptation (LoRA) for domain-specific image generation. The main idea of the proposed system is to combine the text encoding of CLIP, the latent compression of Variational Autoencoder (VAE), and the denoising ability of diffusion to create images that are both semantically relevant and visually coherent based on text prompts. The proposed approach was tested on a Pokemon image-caption dataset for fine-tuning the pre-trained Stable Diffusion model and its effectiveness evaluated. The study shows that the diffusion-based architectures outperform the traditional GAN based methods in terms of image quality, training stability, semantic alignment, and output diversity. The main advantage of LoRA fine-tuning was the substantial decrease in computational load, which involved updating just a small fraction of trainable parameters without compromising the model's performance. Experimental results indicated that successful images of Pokemon could be generated, and that the images were consistent with the text attributes such as color, type, and appearance. The results demonstrate that SD+LoRA is an efficient and scalable domain-specific text-to-image generation system. The research underscores the rising significance of diffusion-based generative AI in digital content creation, imaginative design, entertainment, and cleverness in visual generation systems

Read PDF