Skip to content
Preprint

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

Aug 2026 · 0 citations · 34 references
Physics Computer Science

TL;DR

The first application of sparse-autoencoder-based mechanistic interpretability to particle physics suggests that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.

Abstract

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.

View source

Similar papers

Open access Sep 2026

Toward unified and accurate description of multidimensional nuclear observables via a prior-informed large language model

A prior-informed large language model (LLM) driven multi-task learning framework is proposed for the unified and accurate description of multiple nuclear observables. By fine-tuning the pre-trained DeepSeek-R1-1.5B model with Low-Rank Adaptation (LoRA), lightweight adapters are introduced while preserving general pre...

Shi-Jie Guo, Shou-Yu Wang, E.-H. Wang et al. · 0 citations
Preprint Sep 2026

PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders

PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state, is presented.

Nandita N. Patil, A. EshwarR, G. Honnavar · 0 citations
Preprint Aug 2026

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition, and the spectrum is recoverable from calibration samples at rate $M^{-1/2}$ up to permutation.

Ashim Dhor, Pin-Yu Chen · 0 citations
Open access Sep 2026

Mixture density networks for neutrino reconstruction at hadron colliders

Neutrino momentum reconstruction at hadron colliders is intrinsically ambiguous because the longitudinal momentum is not directly observed. We study this problem in semileptonic $t\bar{t}$ events using \monster{} (Mixture of Neutrino Solutions with Transformer Event Representation), a mixture density network that predi...

Seung-Jin Yang, J. Lee, Junghwan Goh · 0 citations
Preprint Aug 2026

S-matrix informed neural networks for amplitude analysis

This work introduces S-matrix informed neural networks (SINNs), and demonstrates their ability to learn scattering amplitudes directly from data while respecting first principles, and develops a novel data selection procedure which uses the response of constrained neural network ensembles to identify a set of experimen...

W. Smith, A. Rodas, Marius D. Thomas et al. · 0 citations
Preprint Aug 2026

Uncovering Hidden Leptonic Correlations with Flow Matching and Autoencoders

We perform a global search for values of the Yukawa matrices and Majorana masses in the Type-I seesaw mechanism. Using flow matching, which is a generative artificial intelligence (generative AI) method, we generate a broad set of solutions reproducing the experimentally measured values of the neutrino mass-squared dif...

Haruto Kitagawa, Satsuki Nishimura, Hajime Otsuka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.