Skip to content

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Sep 2026 · 0 citations · 74 references
Computer Science

TL;DR

Paint-Anything is presented, which learns a shared hex-prompt interface for generation and editing through object-level color supervision, and introduces Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks.

Abstract

Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics. We present Paint-Anything, which learns a shared hex-prompt interface for generation and editing through object-level color supervision. We develop a data pipeline that constructs Paint-500K from real images through object grounding, perceptual color labeling, and editing-pair synthesis. Since shadows make real-image labels only approximate colors, we complement this supervision with pure-color anchors whose pixels exactly match their paired hex values. These anchors are used only at high-noise timesteps, leaving low-noise training to natural images. We further introduce Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks. On FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3%, respectively, relative to the base model, with ablations supporting the training recipe. It also achieves the highest average CompColor score among the compared methods.

View source

Similar papers

FlowLess: Controlling Abstract Image Generation

A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.

Amir Hertz, Noah Snavely, Google DeepMind · 0 citations
Preprint Sep 2026

Overpainting: Localized Context-aware Diffusion Image Editing

This work implements overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images.

Sam Sartor, Iliyan Georgiev, Michael Fischer et al. · 1 citation
Preprint Sep 2026

GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing

This work introduces GenScale, a benchmark and evaluation protocol for real-world relative object scale in image generation and editing, and introduces Rescale, a model-agnostic post-processing agent for localized scale correction without modifying the source generator.

Ling-Xiao Li, Max Whitton, Ledell Yu Wu et al. · 1 citation · ⚡1
Preprint Sep 2026

Reparametrizing 3D Gaussian Splatting for Real-Time Palette-based Color and Luminance Editing

Professional color editing requires precise control over both color (hue and saturation) and lightness, ideally through separate, independent controls. We present a real-time interactive color editing framework for 3D Gaussian Splatting that supports palette-based recoloring, per-palette tone curves for color-aware lum...

Cheng-Kang Ted Chao, Y. Gingold · 1 citation
Preprint Aug 2026

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

An automated pipeline leveraging Multimodal Large Language Models (MLLMs) is developed to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images, which uniquely enables collaborative spatial-semantic learning.

Weixin Ye, Wei Wang, Hong-Guang Zhu et al. · 0 citations
#computer vision Preprint Sep 2026

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control...

Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.