Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of...