Preprint
Sep 2026
PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders
PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state, is presented.
Nandita N. Patil, A. EshwarR, G. Honnavar
· 0 citations