Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
The first application of sparse-autoencoder-based mechanistic interpretability to particle physics suggests that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.