Skip to content
Open access

CITA-Net: A Unified Multimodal Pipeline for Identity-Preserving Face Editing, Enhanced Captioning, and Text-Based Retrieval

2026 · IEEE Access · Vol 14, pp. 114661-114674 · 0 citations · 36 references
Computer Science

TL;DR

This work introduces CITA-Net, a novel identity-aware face editing architecture based on Stable Diffusion XL, and establishes a strong foundation for prompt-driven face understanding and enables practical applications such as forensic retrieval, semantic face editing, and natural language-based image search.

Abstract

Text-driven face understanding fundamentally depends on the quality of semantic descriptions; however, most existing face datasets provide only coarse or generic captions, limiting both retrieval accuracy and controllable face synthesis. Despite recent advances in multimodal generative models, face-centric applications continue to suffer from two critical challenges: reliable image retrieval from textual descriptions and identity-preserving semantic editing. In this work, we propose a modular yet unified framework that jointly addresses these challenges through enhanced semantic grounding and controlled generative modeling. First, we leverage LLaVA to automatically transform basic annotations in the dataset into rich, fine-grained facial descriptions through a multimodal semantic distillation process, significantly strengthening text–image alignment and improving large-scale text-based face retrieval. Building upon this semantic foundation, we introduce CITA-Net, a novel identity-aware face editing architecture based on Stable Diffusion XL (SDXL). CITA-Net employs an attribute-inversion data construction strategy derived from CelebAMask-HQ, enabling precise semantic disentanglement and controllable facial edits while preserving subject identity. Extensive experiments demonstrate that CITA-Net achieves a superior balance between editability, identity preservation, and visual fidelity, outperforming competitive diffusion-based baselines. In particular, our model attains lower identity loss, enhanced semantic alignment, and improved image quality. By unifying enhanced retrieval with photorealistic, identity-preserving face editing within a single framework, this work establishes a strong foundation for prompt-driven face understanding and enables practical applications such as forensic retrieval, semantic face editing, and natural language-based image search.

Read PDF

Similar papers

Sep 2026

Mitigating Textual Noise in Multimodal FGVC via Hierarchical Semantic Purification and Multi-Stage Alignment

Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous text...

Meng-Huan Zhang, Qing Cai, Fan Zhang et al. · 0 citations
Jul 2026

Sketch-text-driven rectified flow for identity-preserving 4D face generation

Sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation, and a sketch encoder enhanced by Geometric Contour and Texture Detail preprocessing and MixStyle domain adaptation are proposed.

Baodong Wang, Fang Liu, Wei Cao et al. · 0 citations
Open access 2026

Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval

Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...

Zi-Xuan Zhang, Lu-Ming Xiao · 0 citations
Conference 2026

Dual Associations Semantic Enhancement for Image-Text Matching

Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.

Xinlin Zhao · 0 citations
Preprint Aug 2026

SCOUT: Semantic Concept Discovery for Open-Vocabulary Editing of face Recognition Templates

Face recognition templates are compact identity representations, yet they also encode rich semantic information about facial appearance. Prior work has shown that templates can be inverted to images or indirectly manipulated through image-editing pipelines, but direct semantic editing in template space remains largely...

Leon Todorov, Peter Rot, Peter Peer et al. · 0 citations
Aug 2026

PHA-Net: Prototype-based hierarchical alignment network for text-video retrieval

A new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities and introduces multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment.

Xiaolun Jing, Kezhao Yin, Xin-Xing Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.