Artistic Image Generation with Visual Balance Constraints
Abstract
The development of Text-to-Image (T2I) models has made artistic image generation increasingly accessible. However, current diffusion models (e.g., Stable Diffusion XL) still struggle to maintain visual balance and preserve the desired artistic style. Moreover, previous works are often limited to the domain of classical art and heavily rely on users’ prompt-engineering expertise. To address these issues, we propose a novel framework that: (i) leverages Large Language Models (LLMs) to facilitate layout reasoning and prompt decomposition through a coarse-to-fine generation strategy, and (ii) incorporates style-aligned shared attention mechanisms to preserve artistic style during generation. Through both qualitative and quantitative evaluations, our method demonstrates improved visual balance performance while effectively preserving artistic style compared to baseline methods.