Skip to content
Preprint

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

Aug 2026 · 0 citations
Computer Science

TL;DR

An automated pipeline leveraging Multimodal Large Language Models (MLLMs) is developed to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images, which uniquely enables collaborative spatial-semantic learning.

Abstract

Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce **SI-Data**, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called **SI-Edit** that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the [project page](https://github.com/ywxsuperstar/SIEdit).

View source

Similar papers

Aug 2026

Pose-Star++: Semantic-Visual Understanding for Fine-Grained Fashion Image Editing.

Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are...

Yuran Dong, Bo Du, Mang Ye · 0 citations

FlowLess: Controlling Abstract Image Generation

A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.

Amir Hertz, Noah Snavely, Google DeepMind · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Sep 2026

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

IAB edited achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics and shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions.

Chiranjeev Chiranjeev, Muskan Dosi, M. Vatsa et al. · 0 citations
Preprint Sep 2026

SR-Edit: Region-Aware Image Editing via Self-Refinement

SR-Edit is proposed, an image editing framework that overcomes issues via iterative self-refinement and achieves superior preservation and overall image quality compared to existing editing techniques.

Andong Wang, Zehua Chen, Yuxuan Jiang et al. · 0 citations
Preprint Aug 2026

ES3D: Embedding Semantics into 3D Space for Component-Aware Editing

Existing 3D editing methods have made notable progress in controllability, yet they remain limited in several important ways. Most approaches rely on text-driven editing, which struggles to express fine-grained visual changes intended by the user. Moreover, many methods require manually supplied 3D masks or introduce u...

Xuancheng Jin, Ren-Gan Xie, Jiayuan Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.