This work introduces CITA-Net, a novel identity-aware face editing architecture based on Stable Diffusion XL, and establishes a strong foundation for prompt-driven face understanding and enables practical applications such as forensic retrieval, semantic face editing, and natural language-based image search.
Abstract
Text-driven face understanding fundamentally depends on the quality of semantic descriptions; however, most existing face datasets provide only coarse or generic captions, limiting both retrieval accuracy and controllable face synthesis. Despite recent advances in multimodal generative models, face-centric applications continue to suffer from two critical challenges: reliable image retrieval from textual descriptions and identity-preserving semantic editing. In this work, we propose a modular yet unified framework that jointly addresses these challenges through enhanced semantic grounding and controlled generative modeling. First, we leverage LLaVA to automatically transform basic annotations in the dataset into rich, fine-grained facial descriptions through a multimodal semantic distillation process, significantly strengthening text–image alignment and improving large-scale text-based face retrieval. Building upon this semantic foundation, we introduce CITA-Net, a novel identity-aware face editing architecture based on Stable Diffusion XL (SDXL). CITA-Net employs an attribute-inversion data construction strategy derived from CelebAMask-HQ, enabling precise semantic disentanglement and controllable facial edits while preserving subject identity. Extensive experiments demonstrate that CITA-Net achieves a superior balance between editability, identity preservation, and visual fidelity, outperforming competitive diffusion-based baselines. In particular, our model attains lower identity loss, enhanced semantic alignment, and improved image quality. By unifying enhanced retrieval with photorealistic, identity-preserving face editing within a single framework, this work establishes a strong foundation for prompt-driven face understanding and enables practical applications such as forensic retrieval, semantic face editing, and natural language-based image search.
Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous text...
Meng-Huan Zhang, Qing Cai, Fan Zhang et al.· IEEE Transactions on Image P...· 0 citations
Sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation, and a sketch encoder enhanced by Geometric Contour and Texture Detail preprocessing and MixStyle domain adaptation are proposed.
Baodong Wang, Fang Liu, Wei Cao et al.· The Visual Computer· 0 citations
Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...
Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.
Xinlin Zhao· Poster Volume 0008 The 2026...· 0 citations
Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely...
Leon Todorov, Peter Rot, Peter Peer et al.· 0 citations
A new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities and introduces multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment.
Xiaolun Jing, Kezhao Yin, Xin-Xing Yang et al.· Neurocomputing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.