A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control...
Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al.· 0 citations
Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are...
Yuran Dong, Bo Du, Mang Ye· IEEE Transactions on Pattern...· 0 citations
An automated pipeline leveraging Multimodal Large Language Models (MLLMs) is developed to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images, which uniquely enables collaborative spatial-semantic learning.
Weixin Ye, Wei Wang, Hong-Guang Zhu et al.· 0 citations
This work analyzes the semantic structure of abstract art through large-scale embedding visualization, uncovering how perceptual relationships organize artistic meaning, and establishes benchmark tasks for classification, cross-modal retrieval, and text-to-image generation to evaluate how AI models perceive and reprodu...
Hao-Wei Zhang, Yuanpei Zhao, Ji-Zhe Zhou et al.· 0 citations
Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.
Ji Xie, Dewei Zhou, Xin-Yu Huang et al.· 0 citations
Artistic style transfer, which renders content images in the visual style of a reference artwork while preserving semantic structure, is a fundamental problem in computer vision and AIGC, with broad applications in digital art, commercial design, and creative media. Despite progress driven by diffusion-based approaches...