Skip to content

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

Jul 2026 · arXiv.org · Vol abs/2607.24645 · 0 citations · 33 references
Computer Science

TL;DR

FEGA is introduced, an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes, showing that a feature can be interpretable and causally relevant without providing a stable direction for steering.

Abstract

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Enhancing SAE-based Steering via Neighbor Integrated Feature Selection

This paper proposesNeighbor Integrated Feature Selection (NIFS), a plug-and-play strategy that leverages representation similarity to improve feature selection for steering and demonstrates consistent performance gains over conventional top-$k$ selection.

Yutang Liu, Xu Wang, Di-Fan Zou · 1 citation
Preprint Aug 2026

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak · 0 citations
#machine learning Preprint Sep 2026

Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs...

Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder et al. · 0 citations
#machine learning Preprint Sep 2026

Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery

Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is therefore becoming increasingly important for continual model evolution. Sparse autoencoders...

Hendrik Droste, C. M. Adriano, Kathrin Korte et al. · 0 citations
Conference Open access Sep 2026

Learning Local Feature Masks with Variational Information Bottleneck

Instance-wise feature selection (IWFS) identifies informative features for each instance, improving generalization by discarding irrelevant information and enhancing interpretability through personalized explanations. Most IWFS methods adopt a selector--predictor architecture, where a selector generates instance-specif...

Lu Sun, Jun Sakuma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.