The Comprehensive Investigation of Controllable Image Generation Techniques
Abstract
In recent years, the progress of diffusion models in text-to-image generation has been obvious to all. Text-based descriptions alone, however, are often insufficient for precisely controlling screen content. In order to solve this problem, researchers began to try to introduce additional conditions or visual references outside of the text to guide the generation process, which gave birth to the direction of controllable image generation. This article systematically summarizes the latest research results in this field and roughly divides the existing methods into two major categories. The first category is condition-based controllable generation methods, including structural control methods represented by ControlNet, lightweight adaptation solutions such as T2I-Adapter, open set ground truth image generation methods represented by GLIGEN, and multi-task unified frameworks such as UniControl; the second category is example-based methods, covering concept personalized learning methods such as text inversion and DreamBooth, as well as multi-concept combination strategies such as Custom Diffusion. Finally, this article also lists commonly used data sets and evaluation indicators, and comparatively analyzes the performance characteristics of various methods.