Skip to content
Preprint

Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

This work uses simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models, exposing face embeddings as semantically and visually rich biometric representations for web-scale foundation models.

Abstract

Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These models are typically trained for identity discrimination, producing embeddings that are highly effective for biometric matching but largely opaque to semantic interpretation. In contrast, foundation models, pretrained on broad visual or vision--language tasks, provide rich interfaces for describing, retrieving, generating, and organizing visual content. This contrast raises a natural question: what capabilities become available when face embeddings from domain-specific FR models are made interoperable with foundation models? Building on recent work on embedding compatibility across models, we use simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models. Once aligned with a foundation model, a face embedding can be'unmasked'in multiple ways, without training or modifying either model: it can be read in natural language, enabling free-form text queries over a gallery of FR embeddings; rendered into a face image that recovers a person's appearance, using an unmodified diffusion decoder; and converted to a name, enabling identification even in the absence of an enrolled face gallery. In effect, one linear transformation turns an identity embedding into a rich embedding for web-scale foundation models. This interoperability exposes face embeddings as semantically and visually rich biometric representations, with direct implications for interpretability, retrieval, reconstruction, and template security.

View source

Similar papers

Preprint Aug 2026

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, agai...

Guray Ozgur, Mustafa Efe Tamyapar, N. Damer et al. · 0 citations
Preprint Aug 2026

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess representations through discriminative tasks or geometric criteria centered on separability in embedding space. However, strong performance on such evaluations does not es...

Yun Li, Biao Yang, Pei-Xi Wu et al. · 0 citations
#machine learning Preprint Sep 2026

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such a...

Shuang Liang, Le-Jun Liao, Shi-Yuan Zhang et al. · 0 citations
Preprint Aug 2026

Identity-Conditioned Latent Consistency Distillation for Face Synthesis

This work shows that identity-conditioned face synthesis can be performed at a substantially lower computational cost by a latent Consistency Model with few iterations, without compromising image quality for large-scale synthetic face generation.

Tiago Kienen Chaves, Bernardo Biesseck, David Menotti · 0 citations
Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level ma...

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. Ho...

Laurent Colbois, Sébastien Marcel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.