Skip to content
Open access

CITA-Net: A Unified Multimodal Pipeline for Identity-Preserving Face Editing, Enhanced Captioning, and Text-Based Retrieval

2026 · IEEE Access · Vol 14, pp. 114661-114674 · 0 citations · 36 references

Abstract

Text-driven face understanding fundamentally depends on the quality of semantic descriptions; however, most existing face datasets provide only coarse or generic captions, limiting both retrieval accuracy and controllable face synthesis. Despite recent advances in multimodal generative models, face-centric applications continue to suffer from two critical challenges: reliable image retrieval from textual descriptions and identity-preserving semantic editing. In this work, we propose a modular yet unified framework that jointly addresses these challenges through enhanced semantic grounding and controlled generative modeling. First, we leverage LLaVA to automatically transform basic annotations in the dataset into rich, fine-grained facial descriptions through a multimodal semantic distillation process, significantly strengthening text–image alignment and improving large-scale text-based face retrieval. Building upon this semantic foundation, we introduce CITA-Net, a novel identity-aware face editing architecture based on Stable Diffusion XL (SDXL). CITA-Net employs an attribute-inversion data construction strategy derived from CelebAMask-HQ, enabling precise semantic disentanglement and controllable facial edits while preserving subject identity. Extensive experiments demonstrate that CITA-Net achieves a superior balance between editability, identity preservation, and visual fidelity, outperforming competitive diffusion-based baselines. In particular, our model attains lower identity loss, enhanced semantic alignment, and improved image quality. By unifying enhanced retrieval with photorealistic, identity-preserving face editing within a single framework, this work establishes a strong foundation for prompt-driven face understanding and enables practical applications such as forensic retrieval, semantic face editing, and natural language-based image search.

Read PDF