This work proposes CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds by attaching a lightweight Fine-grained Consistency LoRA and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target.
Abstract
Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subject-specific details. A major reason is the lack of high-quality fine-grained identity supervision: real paired data are expensive to collect, while synthesized training pairs often preserve only coarse subject appearance and fail to capture subtle subject-specific details. In this work, we propose CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds. CopyCat performs a one-time refinement of a pretrained subject-to-image model by attaching a lightweight Fine-grained Consistency LoRA (FCLoRA) and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target. This exact self-reconstruction objective substantially simplifies the optimization task, enabling effective fine-grained refinement within only a few seconds. The refinement is performed only once; the resulting model can be directly applied to diverse unseen reference subjects and prompts without further subject-specific optimization. We further revisit subject-to-image LoRA training in double-stream diffusion transformers and find that adapting only the visual stream consistently improves subject consistency. Extensive experiments on DreamBench and XVerseBench demonstrate consistent improvements in fine-grained subject consistency across representative subject-to-image models under both single- and multi-subject settings.
Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (M...
Jihun Park, Kyoungmin Lee, Jongmin Gim et al.· 0 citations
Commercial and advertising images are frequently affected by poor framing, partially cropped subjects, truncated text or logos, and insufficient context, all of which can reduce subject clarity, i.e., the ability of an image to clearly communicate its primary subject. Image outpainting offers a scalable solution by ext...
Abhilash Neog, Taewan Kim, Yi Wu et al.· 0 citations
PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.
Ling-Chen Sun, Rong-Yuan Wu, Xiang-Tao Kong et al.· 1 citation
The RenderMatte dataset is constructed, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets that features exact strand-level alpha annotations and diverse background composites, demonstrating a scalable path toward high-fidelity matting in open-world scenes.
Ze-Cheng Ren, Ya-Fei Hu, Jianing Zhao et al.· 0 citations
PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.
Yu-Feng Chi, Hui-Min Ma, Fan Gao et al.· 0 citations
FIT employs a lightweight Degradation Encoder to predict a global degradation vector and a spatial degradation map from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation, and introduces a task-token dropout strategy that regularizes task conditioning du...
Zi-Hao He, Yunfeng Wu, Xinchao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.