This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.
Abstract
Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this transfer works: a face-verification protocol measures it, and an ablation changing only the adapter isolates its contribution. Not every concept survives, because an FR model earns its invariances by discarding the factors it must verify identities across. A label-free detectability measure compares each concept's separability in FR space against the VLM space, and the 100 most detectable form the model's readable semantic signature, which separates identities better than the full vocabulary. We cover four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations. We benchmark attribute-level auditing under three supervision settings, human labels (current practice), VLM pseudo-labels, and our fully prompt-driven audit, against real verification behavior. With no labels, the prompt-driven audit ranks four FR models by their measured per-ethnicity RFW errors and ranks controlled attribute changes by their true verification cost.
This work uses simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models, exposing face embeddings as semantically and visually rich biometric representations for web-scale foundation models.
Despite-encoder vision-language models expose a similarity interface that enables zero-shot retrieval but fails compositional constraints, this work proposes factored inference, which separates evidence extraction from constraint execution, and introduces LCSE (Logic-Constrained Score Editing), a training-free method t...
S. Alshehri, Zhan-Tao Yang, Han Zhang et al.· arXiv.org· 0 citations
Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evidence that enables users to understand, assess, and challenge their decisions. Vision-language models (VLMs) provide a promising foundation for explainable face recogniti...
A. Estrada-Real, Lydia Alapatt, C. Busch et al.· 0 citations
These results establish frozen DINOv3 as a strong zero-shot representation for region-level facial correspondence and identify intermediate self-supervised features as the most useful layer for dense face analysis.
Izaldein Al-Zyoud, Abdulmotaleb El Saddik· arXiv.org· 0 citations
The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference, and that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.
PatchGate is proposed, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions.
Jihyung Ko, Eunji Jung, Hyeongsub Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.