ExplorerArtificial IntelligenceAI
Research PaperResearchia:202608.27046

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

Raphaël Bonnet-Guerrini

Abstract

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head bar...

Submitted: August 27, 2026Subjects: AI; Artificial Intelligence

Description / Details

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At 20%20\% selection efficiency, this interpretable estimator improves the median angular resolution from 20.220.2^\circ to 3.23.2^\circ. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.


Source: arXiv:2608.26090v1 - http://arxiv.org/abs/2608.26090v1 PDF: https://arxiv.org/pdf/2608.26090v1 Original Link: http://arxiv.org/abs/2608.26090v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Aug 27, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark