Practical Consequences of FID–CLIP Disagreement in Low-Data Diffusion Model Fine-Tuning
Diffusion models are increasingly fine-tuned for domain-specific image generation, yet fine-tuning strategies are usually selected with a single evaluation metric. This paper examines when the Fréchet Inception Distance (FID) and a text–image similarity score based on Contrastive Language–Image Pre-training (CLIP) disagree in the ranking of such strategies. The study evaluates 68 configurations built on Stable Diffusion 1.5 with low-rank adaptation: on each of four datasets that span style transfer and subject-driven personalization, a clustering-based curriculum and a random-order baseline are matched across five data-availability regimes, plus seven shared ablation and baseline controls. Three results stand out. The FID–CLIP relationship changes in character across datasets, from a strongly positive association to a sign reversal between the rank and linear correlations. The two metrics select different winners in most head-to-head comparisons, and the practical cost of following the wrong metric ranges from negligible on style transfer to severe on personalization. The random-order baselines are FID-optimal in most comparisons, training time tends to agree with the FID-optimal choice, and a lower training loss is a poor proxy for generation quality. Because each configuration is evaluated with a small generated sample of 30 to 40 images, FID is read throughout as a comparative diagnostic under a fixed protocol and its absolute values are not interpreted. Robustness checks with larger evaluation sets and multi-seed reruns confirm the large-gap conclusions, whereas the close-call cases prove fragile. A validation on Stable Diffusion XL gives initial evidence that these patterns are not specific to one architecture.