Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
FEGA is introduced, an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes, showing that a feature can be interpretable and causally relevant without providing a stable direction for steering.