FEGA is introduced, an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes, showing that a feature can be interpretable and causally relevant without providing a stable direction for steering.
Abstract
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.
This paper proposesNeighbor Integrated Feature Selection (NIFS), a plug-and-play strategy that leverages representation similarity to improve feature selection for steering and demonstrates consistent performance gains over conventional top-$k$ selection.
It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak· 0 citations
The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs...
Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder et al.· 0 citations
Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is therefore becoming increasingly important for continual model evolution. Sparse autoencoders...
Hendrik Droste, C. M. Adriano, Kathrin Korte et al.· 0 citations
Instance-wise feature selection (IWFS) identifies informative features for each instance, improving generalization by discarding irrelevant information and enhancing interpretability through personalized explanations. Most IWFS methods adopt a selector--predictor architecture, where a selector generates instance-specif...
Lu Sun, Jun Sakuma· Proceedings of the Thirty-Fi...· 0 citations
A systematic study of how pruning affects SAE behavior is presented and theoretically shows that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm.
Suchit Gupte, Xue-Ru Zhang, M. Khalili· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.