An automated pipeline leveraging Multimodal Large Language Models (MLLMs) is developed to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images, which uniquely enables collaborative spatial-semantic learning.
Abstract
Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce **SI-Data**, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called **SI-Edit** that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the [project page](https://github.com/ywxsuperstar/SIEdit).
Fashion image editing demands high-dimensional, fine-grained control to follow personalized, unpredictable natural-language instructions. Yet current methods are limited by a fundamental trade-off: fashion-specific approaches offer structural accuracy but lack semantic flexibility, while general text-driven editors are...
Yuran Dong, Bo Du, Mang Ye· IEEE Transactions on Pattern...· 0 citations
A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.
Amir Hertz, Noah Snavely, Google DeepMind· 0 citations
Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...
Trong-Tai Dam Vu, Vinh-Tiep Nguyen· International Conference on...· 0 citations
IAB edited achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics and shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions.
Chiranjeev Chiranjeev, Muskan Dosi, M. Vatsa et al.· 0 citations
SR-Edit is proposed, an image editing framework that overcomes issues via iterative self-refinement and achieves superior preservation and overall image quality compared to existing editing techniques.
Andong Wang, Zehua Chen, Yuxuan Jiang et al.· 0 citations
Existing 3D editing methods have made notable progress in controllability, yet they remain limited in several important ways. Most approaches rely on text-driven editing, which struggles to express fine-grained visual changes intended by the user. Moreover, many methods require manually supplied 3D masks or introduce u...
Xuancheng Jin, Ren-Gan Xie, Jiayuan Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.