This work explores the latent space of a pre-trained, frozen encoder for text-to-image personalization, and shows that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits.
Abstract
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity. To enable fine-grained latent identity tuning, we explore the latent space of a pre-trained, frozen encoder for text-to-image personalization. Our approach requires no additional training. Instead, it leverages the existing architecture of a frozen encoder to uncover latent semantic directions. This space consists of a set of latent tokens that play distinct roles in capturing different aspects of an identity and often correspond to specific spatial or semantic facial regions. We show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits. We validate our approach through qualitative and quantitative experiments that demonstrate diverse localized facial edits while preserving cross-image identity consistency. Project page at: https://garibida.github.io/IdentityTuning/
This work shows that identity-conditioned face synthesis can be performed at a substantially lower computational cost by a latent Consistency Model with few iterations, without compromising image quality for large-scale synthetic face generation.
Tiago Kienen Chaves, Bernardo Biesseck, David Menotti· 0 citations
Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surroundi...
Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for...
J. Lee, Yujin Kim, Ghazanfar Ali et al.· 0 citations
Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...
Kuang-Rong Hao· International Conference on...· 0 citations
Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...