Skip to content
Preprint

CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work proposes CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds by attaching a lightweight Fine-grained Consistency LoRA and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target.

Abstract

Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subject-specific details. A major reason is the lack of high-quality fine-grained identity supervision: real paired data are expensive to collect, while synthesized training pairs often preserve only coarse subject appearance and fail to capture subtle subject-specific details. In this work, we propose CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds. CopyCat performs a one-time refinement of a pretrained subject-to-image model by attaching a lightweight Fine-grained Consistency LoRA (FCLoRA) and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target. This exact self-reconstruction objective substantially simplifies the optimization task, enabling effective fine-grained refinement within only a few seconds. The refinement is performed only once; the resulting model can be directly applied to diverse unseen reference subjects and prompts without further subject-specific optimization. We further revisit subject-to-image LoRA training in double-stream diffusion transformers and find that adapting only the visual stream consistently improves subject consistency. Extensive experiments on DreamBench and XVerseBench demonstrate consistent improvements in fine-grained subject consistency across representative subject-to-image models under both single- and multi-subject settings.

View source

Similar papers

Preprint Aug 2026

CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets

Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (M...

Jihun Park, Kyoungmin Lee, Jongmin Gim et al. · 0 citations
Preprint Sep 2026

Preserving Subject-Clarity in Image Outpainting with Multiscale Wavelet Supervision

Commercial and advertising images are frequently affected by poor framing, partially cropped subjects, truncated text or logos, and insufficient context, all of which can reduce subject clarity, i.e., the ability of an image to clearly communicate its primary subject. Image outpainting offers a scalable solution by ext...

Abhilash Neog, Taewan Kim, Yi Wu et al. · 0 citations
Preprint Aug 2026

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

PixRestore is presented, a VAE-free pixel-space Diffusion Transformer (DiT) for UIR, where the diffusion backbone is trained entirely from scratch, without relying on T2I pretraining.

Ling-Chen Sun, Rong-Yuan Wu, Xiang-Tao Kong et al. · 1 citation
Preprint Aug 2026

RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting

The RenderMatte dataset is constructed, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets that features exact strand-level alpha annotations and diverse background composites, demonstrating a scalable path toward high-fidelity matting in open-world scenes.

Ze-Cheng Ren, Ya-Fei Hu, Jianing Zhao et al. · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
Preprint Aug 2026

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

FIT employs a lightweight Degradation Encoder to predict a global degradation vector and a spatial degradation map from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation, and introduces a task-token dropout strategy that regularizes task conditioning du...

Zi-Hao He, Yunfeng Wu, Xinchao Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.