Skip to content
Conference Open access

Enhanced Text-to-Image Editing with Multi-Step Control and Explainability

2025 · Proceedings of the 1st International Conference on Interdisciplinary Research in Science, Engineering, and Technology · pp. 113-123 · 0 citations · 17 references

TL;DR

This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary.

Abstract

: Text-driven image editing has advanced significantly in generating and modifying visual content. Existing approaches often face challenges in maintaining visual coherence across sequential edits and providing informative rationales for alterations. This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary. Through the combination of robust fine-tuning techniques and leading-edge visual understanding models, the framework enhances edit consistency, image quality, and user control. Qualitative and illustrative quantitative results demonstrate the effectiveness of the InstructPix2Pix-MB-FT model in performing instruction-driven image editing tasks, achieving high realism and fidelity in object modification, scene enhancement, and human feature changes. The developed method has the potential to be used in creative design, content generation, and visual storytelling.

Read PDF

Similar papers

Preprint Oct 2026

When Text-to-Image Helps Editing: The Effects of Conditioning During Denoising

Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pu...

Lidia Troeshestova, A. Ustyuzhanin, Sergey Kastryulin · 0 citations
Preprint Aug 2026

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible...

Yu-Qian Zhou, Zhenghong Zhou, Zongze Wu et al. · 1 citation
Open access Aug 2026

Text-to-hierarchical three-dimensional scene generation: a new approach for layered three-dimensional modeling from natural language

An end-to-end hierarchical framework for text-to-3D scene generation that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction is introduced, offering a powerful solution for applications, such as virtual reality and digital twins.

Zuan Gu, Tian-Han Gao, Lang-Xu Zhao et al. · 0 citations
Preprint Sep 2026

Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering

Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stroke structures and of...

Shu-Lian Zhang, Xiang-Yu Shu, Wen-Bo Li et al. · 0 citations
#machine learning Conference Open access Sep 2026

STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding

We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation...

Nicolas Thiébaut, Nameer Hirschkind, Xiao Yu et al. · 0 citations
Preprint Aug 2026

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.

Hongyi Fang, Chu-Wen Xie, Ben-Jia Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.