Skip to content
Preprint

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

IAB edited achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics and shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions.

Abstract

Instruction-guided image editing has a training-time blind spot. Generative editors are never required to semantically verify whether their outputs actually satisfy the instruction. Supervision stops at reconstruction and input textual-level conditioning. This produces incomplete edits, spatial spillover, and poor localization. We present IABEdit, a model-agnostic framework that embeds differentiable semantic verification into training. A frozen vision-language model extracts spatially-aware descriptors from the ground-truth edit. A trainable aligner then reproduces them from the generated output. The residual between the two becomes a gradient that teaches the generator both what to edit and where, with no inference-time VLM cost. IABEdit is compatible with diverse backbones, including U-Net (Stable Diffusion) and MMDiT (FLUX), without altering their inference pipelines. On MagicBrush, it improves structural fidelity by +3.49 DINO-I over the best diffusion baseline and +1.26 over the best overall baseline, while remaining competitive on instruction alignment. It also achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics. Most consequentially, on the D-LORD surveillance benchmark, it surpasses the proprietary Gemini agent by +5.13 DINO-P under heavy occlusion, where preserving identity is hardest. This shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions. Human and GPT-4o evaluations confirm perceptually precise, well-localized edits.

View source

Similar papers

Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Sep 2026

Learning 3D Editing without Paired Supervision via Generative Prior Distillation

Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottleneck: the severe scarcity of high-quality paired training data. Existing approaches attempt to bypass this by either relying on slow test-time optimization or training on pseudo-pairs constructed via complex pi...

Hao Wen, Wei-Bin Yun, Hongxing Fan et al. · 0 citations
Preprint Aug 2026

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editi...

Zhe-Fan Rao, Bin-Yi Zou, Xuanhua He et al. · 0 citations
Preprint Sep 2026

Attention-Scoped Guidance: Training-Free Spatial Control for Image Editing

Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit and one that pulls back toward the source image, using global...

Ze-Yan Li, Wei Zhou, Hadi Amirpour et al. · 0 citations
Preprint Sep 2026

DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling

Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matchi...

Zhe-Yuan Zhan, Can Wang, Jia-Wei Chen et al. · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.