Skip to content
Preprint

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

Aug 2026 · 0 citations · 87 references
Computer Science

TL;DR

The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition, and the spectrum is recoverable from calibration samples at rate $M^{-1/2}$ up to permutation.

Abstract

Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.

View source

Similar papers

#machine learning Preprint Sep 2026

PhysSAE: Mechanistic Interpretability with Sparse Autoencoders

Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framewor...

Nandita N. Patil, Eshwar R. A., Gajanan V. Honnavar · 0 citations
Preprint Aug 2026

A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource

This article presents the abridged core of \emph{A Mathematical Theory of Interpretation} (MTI), which treats interpretation as observer-relative spectral measurement under an access structure. MTI makes interpretation a method-design problem: access, query, utility, and medium determine what an observer can select, id...

B. Reynolds · 0 citations
Preprint Aug 2026

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

The first application of sparse-autoencoder-based mechanistic interpretability to particle physics suggests that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.

Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, I. Timiryasov et al. · 0 citations
Preprint Aug 2026

A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation

An idealized model where a conditional computation is carried additively through a residual stream, F(x)=F_0(x)+\sum_i\alpha_i(x)v_i$, read out by a linear functional is studied, and three exact results are proved, including an exact first-order interaction formula with a provably second-order remainder.

Abdallah Khemais · 1 citation
#artificial intelligence Preprint Aug 2026

Efficient Auto-Interpretability of AI Models in Biology

Cross-seed dictionary stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations.

Piotr Jedryszek, O. Crook · 0 citations
#machine learning Preprint Aug 2026

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

The method proposed leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) pred...

Tom A. Splittgerber, Niklas Koenen, Marvin N. Wright et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.