Skip to content

Latent-Identity Tuning in Text-to-Image Personalization Models

Jul 2026 · arXiv.org · Vol abs/2607.11885 · 0 citations · 77 references
Computer Science

TL;DR

This work explores the latent space of a pre-trained, frozen encoder for text-to-image personalization, and shows that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits.

Abstract

Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity. To enable fine-grained latent identity tuning, we explore the latent space of a pre-trained, frozen encoder for text-to-image personalization. Our approach requires no additional training. Instead, it leverages the existing architecture of a frozen encoder to uncover latent semantic directions. This space consists of a set of latent tokens that play distinct roles in capturing different aspects of an identity and often correspond to specific spatial or semantic facial regions. We show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits. We validate our approach through qualitative and quantitative experiments that demonstrate diverse localized facial edits while preserving cross-image identity consistency. Project page at: https://garibida.github.io/IdentityTuning/

View source

Similar papers

Preprint Aug 2026

Identity-Conditioned Latent Consistency Distillation for Face Synthesis

This work shows that identity-conditioned face synthesis can be performed at a substantially lower computational cost by a latent Consistency Model with few iterations, without compromising image quality for large-scale synthetic face generation.

Tiago Kienen Chaves, Bernardo Biesseck, David Menotti · 0 citations
Preprint Sep 2026

Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surroundi...

Meng-Wei Ren, Xuan-Er Zhang, Zhi-Hao Xia · 0 citations
Preprint Aug 2026

Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model

Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for...

J. Lee, Yujin Kim, Ghazanfar Ali et al. · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations
Open access 2026

Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval

Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity i...

Zi-Xuan Zhang, Lu-Ming Xiao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.