Preprint
Aug 2026
SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable
This work introduces SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification, and extracts interpretable rule-sets using the SAE neurons to create neurosymbolic models.
Philip T. H. Lee, Parth Padalkar
· 0 citations