Preprint
Aug 2026
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
The first application of sparse-autoencoder-based mechanistic interpretability to particle physics suggests that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.
Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, I. Timiryasov et al.
· 0 citations