PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state, is presented.
Abstract
Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: $h_{\mathrm{cf}} = h - \alpha z_k d_k$, bypassing the SAE decoder entirely. Across six PDE families, with 3 PINN seeds and 3 SAE seeds each---we show that (i) Our discovered SAE atoms align with independently-defined physical observables (max Pearson $|r|=0.951$, always $\gg$ permutation null), (ii) the causal footprint of top-aligned atom ablation is 1.2--4.2$\times$ more spatially concentrated canonical than PCA or ICA interventions, and (iii) top-aligned atoms outperform matched random controls on causal localization for structured physical concepts (ESF$_{80}$ advantage 0.04-0.44). Two-atom bilateral representations improve concept regression R$^2$ by $\Delta R^2\!=\!0.05\text{-}0.15$ over single atoms, while random pairs decrease it by up to 0.60. These results demonstrate that PINNs develop sparse, physically structured latent representations that can be identified and causally interrogated post-hoc, opening a path toward interpretability-aware scientific machine learning.
The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition, and the spectrum is recoverable from calibration samples at rate $M^{-1/2}$ up to permutation.
The results suggest that trained SAEs need not reach the corresponding minima, and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.
PINNA is validated across three fundamentally different benchmarks: two composite‐material problems involving nonlinear stress‐strain behavior and multistage failure, and a large‐scale 1‐D laminar combustion problem governed by stiff chemical kinetics, thermal transport, and reduced fluid mechanics.
Zheng-Tao Yao, Philippe Hawi, V. Aitharaju et al.· International Journal for Nu...· 0 citations
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, diction...
Gautam Ranka, S. Pandere, Aiden Dsouza· 0 citations
HiPACE is introduced, an evaluation protocol that tests the boundary's structural consequence in real SAE dictionaries--measuring parent--child decoder structure over WordNet families, freezing the discovery-selected statistic before testing on unseen families, and contrasting genuine families against randomized siblin...
Jin-Yuan Zhang, Peng-Ji He, Yin Yuan et al.· 0 citations
This work introduces SpIn-ViT, a framework that jointly trains a pretrained ViT and a modified SAE end-to-end, directly aligning sparse patch-level representations with image classification, and extracts interpretable rule-sets using the SAE neurons to create neurosymbolic models.
Philip T. H. Lee, Parth Padalkar· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.