Skip to content
Preprint

SR-Edit: Region-Aware Image Editing via Self-Refinement

Sep 2026 · 0 citations · 72 references
Computer Science

TL;DR

SR-Edit is proposed, an image editing framework that overcomes issues via iterative self-refinement and achieves superior preservation and overall image quality compared to existing editing techniques.

Abstract

With the recent rapid progress in generative models, image editing has made remarkable advances, yet achieving faithful edits that precisely modify only the target regions while strictly preserving all other regions remains challenging. Since externally provided region annotations are often difficult to obtain in practice, a growing body of work seeks to improve preservation by automatically inferring edit and non-edit regions, and then enforcing consistency on the latter. However, these approaches still suffer from inaccurate region estimation and heuristic correction strategies that distort the native inference process, making methods designed for fidelity themselves a new source of artifacts. We propose SR-Edit, an image editing framework that overcomes these issues via iterative self-refinement. Specifically, at each iteration, SR-Edit first (i) extracts progressively precise and self-consistent region separation from the model's own predictions by lightweight post-processing, and then (ii) enforces preservation in non-edit areas through correction updates that remain aligned with the original sampling dynamics. Extensive experiments demonstrate that SR-Edit achieves superior preservation and overall image quality compared to existing editing techniques.

View source

Similar papers

#computer vision Preprint Aug 2026

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

The proposed MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions, incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it.

Rui Xu, Yang Yong, Shun-Zi Yang et al. · 0 citations
Preprint Sep 2026

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed...

Yu-Long Chen, Zi-Qian Zhang, Hao-Yu Zhang et al. · 0 citations
Preprint Aug 2026

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

Jiayi Song, Shijie Huang, Fang-Tai Wu et al. · 0 citations
#machine learning Preprint Sep 2026

Constrained Edit Fields for Training-Free Flow Editing

Text-guided image editing aims to perform a desired edit while preserving source content unrelated to it. Pretrained rectified-flow models enable training-free editing of real images through modifications to their sampling trajectories. However, responses at locations unrelated to the desired edit can still accumulate...

Jing-Xuan Kang, Yin-Song Wang, Che Liu et al. · 0 citations
Preprint Sep 2026

Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness

In recent years, diffusion stochastic differential equation (SDE) inversion and inversion-free methods have become prevalent for training-free image editing, as they can achieve faithful reconstruction without tuning. However, existing approaches remain inefficient, exhibit limited plasticity, and struggle to accuratel...

Hai-Yan Wei, Yunlong Wang, Huaibo Huang et al. · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.