Skip to content

How are linear representations learned? Exact solutions to the dynamics of abstraction

Jul 2026 · arXiv.org · Vol abs/2607.08843 · 0 citations
Computer Science

TL;DR

A striking attenuation law is proved: both nonlinearities weaken abstraction in activations relative to preactivations, and the theory is applied to improve linear probe generalization in LLMs.

Abstract

In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call"abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.

View source

Similar papers

Jul 2026

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

VisualPatchWorld is introduced, which represents world dynamics as code and first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error.

Jiaxin Bai, Jia–Jie Xiong · 0 citations
#machine learning Review Sep 2026

Feature Superposition in Neural Networks: From Theory to Practice

Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptio...

Dai Shi, Xiao-Yu Li, Andi Han et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Study of Hidden-State Optimization Order in Predictive Coding Networks

A boundary-first inference schedule that partitions a model into chunks, first coordinates hidden states at chunk boundaries, and then refines representations within each chunk is proposed, which instantiate in predictive coding networks (PCNs), a local-learning framework in which hidden activities and prediction error...

Xue-Yuan Li, Danilo Vasconcellos Vargas · 0 citations
Review Open access Aug 2026

A survey of theory-grounded interpretability for deep neural networks

Deep neural networks achieve state-of-the-art performance across computer vision, natural language processing, and scientific discovery, yet their internal decision-making mechanisms remain difficult to interpret. Post-hoc explanation methods such as SHAP, saliency maps, and attention visualization provide empirical in...

M. U. Khalid, Abeer A. K. Alharbi, Shariq Bashir et al. · 0 citations
Conference Open access Sep 2026

Context-Adaptive World Models for Visual Out-of-Distribution Generalization

This thesis is developing a context-adaptive Mamba-based world model that infers K context vectors from short pixel-trajectory windows, and shows that its suboptimality is bounded by √K times the context-inference error along a Lipschitz chain through FiLM, the SSM, and the decoder.

Shubham Subhnil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.