2025· Proceedings of the 1st International Conference on Interdisciplinary Research in Science, Engineering, and Technology· pp. 113-123· 0 citations· 17 references
TL;DR
This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary.
Abstract
: Text-driven image editing has advanced significantly in generating and modifying visual content. Existing approaches often face challenges in maintaining visual coherence across sequential edits and providing informative rationales for alterations. This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary. Through the combination of robust fine-tuning techniques and leading-edge visual understanding models, the framework enhances edit consistency, image quality, and user control. Qualitative and illustrative quantitative results demonstrate the effectiveness of the InstructPix2Pix-MB-FT model in performing instruction-driven image editing tasks, achieving high realism and fidelity in object modification, scene enhancement, and human feature changes. The developed method has the potential to be used in creative design, content generation, and visual storytelling.
Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pu...
Lidia Troeshestova, A. Ustyuzhanin, Sergey Kastryulin· 0 citations
Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible...
Yu-Qian Zhou, Zhenghong Zhou, Zongze Wu et al.· 1 citation
An end-to-end hierarchical framework for text-to-3D scene generation that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction is introduced, offering a powerful solution for applications, such as virtual reality and digital twins.
Zuan Gu, Tian-Han Gao, Lang-Xu Zhao et al.· Visual Computing for Industr...· 0 citations
Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stroke structures and of...
Shu-Lian Zhang, Xiang-Yu Shu, Wen-Bo Li et al.· 0 citations
We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation...
Nicolas Thiébaut, Nameer Hirschkind, Xiao Yu et al.· Pacific-Asia Conference on K...· 0 citations
This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.
Hongyi Fang, Chu-Wen Xie, Ben-Jia Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.