Skip to content
Preprint

Overpainting: Localized Context-aware Diffusion Image Editing

Sep 2026 · 1 citation · 58 references
Computer Science

TL;DR

This work implements overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images.

Abstract

We present"overpainting", an image editing operation which offers both control over the location of the edit and awareness of the previous content in that location. The overpainted area is given by a trimap, where white-annotated pixels must be edited, gray-annotated pixels may be edited, and black-annotated pixels must not be edited. This enables both precise and loose control, depending on user intent. We implement overpainting by adapting a pretrained image editing diffusion model using a combination of joint attention and low-rank adaption across input images with attention-dropout to balance the information flow between noise, source and mask images. We present a novel, automated, training data generation pipeline that (1) generates a set of candidate image pairs leveraging existing language-based editing models, (2) carefully curates those pairs, and (3) extracts a trimap from each usable pair. We demonstrate the versatility of our overpainting model on a wide range of editing tasks.

View source

Similar papers

Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength

Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotati...

Candi Zheng, Yuan Lan · 0 citations
Preprint Aug 2026

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.

Hongyi Fang, Chu-Wen Xie, Ben-Jia Zhou et al. · 0 citations
#computer vision Preprint Aug 2026

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

The proposed MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions, incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it.

Rui Xu, Yang Yong, Shun-Zi Yang et al. · 0 citations
Preprint Aug 2026

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.

Long Cui, Xiao-Qian Liu, Qi Qin et al. · 0 citations
Preprint Aug 2026

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

Jiayi Song, Shijie Huang, Fang-Tai Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.