The first application of sparse-autoencoder-based mechanistic interpretability to particle physics suggests that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.
Abstract
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.
A prior-informed large language model (LLM) driven multi-task learning framework is proposed for the unified and accurate description of multiple nuclear observables. By fine-tuning the pre-trained DeepSeek-R1-1.5B model with Low-Rank Adaptation (LoRA), lightweight adapters are introduced while preserving general pre...
Shi-Jie Guo, Shou-Yu Wang, E.-H. Wang et al.· Chinese Physics Letters· 0 citations
PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state, is presented.
Nandita N. Patil, A. EshwarR, G. Honnavar· 0 citations
The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition, and the spectrum is recoverable from calibration samples at rate $M^{-1/2}$ up to permutation.
Neutrino momentum reconstruction at hadron colliders is intrinsically ambiguous because the longitudinal momentum is not directly observed. We study this problem in semileptonic $t\bar{t}$ events using \monster{} (Mixture of Neutrino Solutions with Transformer Event Representation), a mixture density network that predi...
This work introduces S-matrix informed neural networks (SINNs), and demonstrates their ability to learn scattering amplitudes directly from data while respecting first principles, and develops a novel data selection procedure which uses the response of constrained neural network ensembles to identify a set of experimen...
W. Smith, A. Rodas, Marius D. Thomas et al.· 0 citations
We perform a global search for values of the Yukawa matrices and Majorana masses in the Type-I seesaw mechanism. Using flow matching, which is a generative artificial intelligence (generative AI) method, we generate a broad set of solutions reproducing the experimentally measured values of the neutrino mass-squared dif...