SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision
An automated pipeline leveraging Multimodal Large Language Models (MLLMs) is developed to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images, which uniquely enables collaborative spatial-semantic learning.