Skip to content

Naming the Concepts Classifiers Rely On: Language-Anchored Decomposition for Faithful Explanation

Jul 2026 · arXiv.org · Vol abs/2607.07264 · 0 citations · 27 references
Computer Science

TL;DR

Across natural-image, scene, and medical-imaging benchmarks, LAD produces spatially precise explanations that are decision-relevant under both concept insertion and deletion, while uniquely providing stable, human-interpretable concept names.

Abstract

Deep neural networks are widely deployed in high-stakes visual applications where interpretability is critical, yet existing explanations face a trade-off: post-hoc concept methods recover factors that are faithful to a model's behavior but unnamed, while naming and by-design methods attach human-readable concepts only by retraining or altering the classifier. We propose Language-Anchored Decomposition (LAD), a post-hoc framework that delivers concepts which are simultaneously named, faithful, and obtained without modifying the model. For each class, a large language model proposes a concept vocabulary that CLIP-based similarity maps localize across image regions. Inverting standard non-negative matrix factorization, LAD fixes these language-grounded maps as the coefficient matrix and learns only a concept basis that reconstructs the frozen encoder's activations, so naming becomes a structural constraint and the model's own feature geometry determines which concepts are retained. Removing this anchor preserves accuracy but collapses attribution faithfulness. Across natural-image, scene, and medical-imaging benchmarks, LAD produces spatially precise explanations that are decision-relevant under both concept insertion and deletion, while uniquely providing stable, human-interpretable concept names.

View source

Similar papers

Jul 2026

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

Despite-encoder vision-language models expose a similarity interface that enables zero-shot retrieval but fails compositional constraints, this work proposes factored inference, which separates evidence extraction from constraint execution, and introduces LCSE (Logic-Constrained Score Editing), a training-free method t...

S. Alshehri, Zhan-Tao Yang, Han Zhang et al. · 0 citations
Jul 2026

What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

PeakPatch is proposed, a lightweight post-hoc correction system that intercepts the CLIP text encoder at its compositional peak and recovers the lost negation signal without altering pretrained weights.

Chen-Yi Lu, Yueh-Shao Chen, S. Chaterji · 0 citations
Preprint Aug 2026

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, agai...

Guray Ozgur, Mustafa Efe Tamyapar, N. Damer et al. · 0 citations
Preprint Aug 2026

Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

This work uses simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models, exposing face embeddings as semantically and visually rich biometric representations for web-scale foundation models.

Fizza Rubab, Yi-Ying Tong, Arun Ross · 0 citations
Jul 2026

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

LeapBot-WA establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor and introduces the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift.

Pei Liu, Nan Zheng, Lang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.