Skip to content
Preprint

SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work introduces SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification, and extracts interpretable rule-sets using the SAE neurons to create neurosymbolic models.

Abstract

Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features. However, because post-hoc SAEs are trained on frozen representations after the ViT has already been optimized, their latent features are not directly aligned with the downstream classification objective. We introduce SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification. SpIn-ViT learns semantically coherent neuron activations that localize meaningful image regions while maintaining competitive predictive performance. We evaluate SpIn-ViT across nine image-classification benchmarks using classification accuracy, quantitative interpretability metrics, AI-based and Human evaluations. Compared with the previous state-of-the-art post-hoc SAE method, SpIn-ViT achieves 8.84% higher average classification accuracy, an AI-based interpretability score nearly four times as high, and a human-evaluation score more than twice as high. We further extract interpretable rule-sets using the SAE neurons to create neurosymbolic models which achieve 5.97% higher average classification accuracy while requiring a 58.8\% smaller rule-set than the neurosymbolic models created from the SOTA post-hoc SAE method.

View source

Similar papers

Open access 2026

HIFN-Transformer: Learnable Information-Theoretic Parameters for Interpretable Deep Classification

HIFN-T is presented, a framework extending the Variational Information Bottleneck through four jointly learnable per-layer parameters: information retention, entropy budget, magnitude scaling, and global information gates that generalizes standard VIB as a special case and characterize the role of the entropy budget as...

Mohammed Tawfik · 0 citations
Preprint Sep 2026

PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders

PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state, is presented.

Nandita N. Patil, A. EshwarR, G. Honnavar · 0 citations

UvA-DARE (Digital Academic Repository) Elastic ViTs from Pretrained Models without Retraining

SnapViT: single-shot network approximation for pruned Vision Transformers is introduced, a new post-pretraining structured pruning method that enables elastic inference across a continuum of compute budgets, and a self-supervised importance scoring mechanism that maintains strong performance without requiring retrainin...

Walter Simoncini, Michael Dorkenwald, Tijmen Blankevoort et al. · 0 citations
Preprint Aug 2026

OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification

Prototypical part-based models provide explainable predictions by comparing input regions to learned prototypes. However, current approaches are burdened by complex, multi-stage training pipelines and heavily rely on auxiliary regularization to prevent prototype collapse. To overcome these limitations, we introduce Ort...

I. Carretero, G. Angulo, R. del Amor et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.