This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which is called feature recall, and defines it, shows it applies across architectures, and contrasts it with the established paradigm of feature combination.
Abstract
Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.
Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with which they accomplish time-series prediction remain unclear. Specifically, whether they tru...
Rahul Chowdhury, Timothy Rupprecht, Senhao Cao et al.· 0 citations
Although performance of language-based models is improved by scaling, whether the gap to a structure-aware architecture can eventually be eliminated remains untested.
How does compositional reasoning emerge during learning? We address this question by mathematically analyzing the learning dynamics of deep linear networks. We train these networks in structured synthetic environments and derive a theory linking the structure of experience to compositional learning. Our theory predicts...
A three-stage symbolic reasoning circuit (abstraction, induction, retrieval) is traced across shot counts in three model families and finds it is detectable and functional well before the model achieves high accuracy.
This tutorial provides a comprehensive, end-to-end view of LLM interpretability, transitioning from microscopic neural analysis to macroscopic application and deployment, and explores how these interpretability paradigms scale and inspire the design of frontier architectures, agentic systems, and thinking models.
Wei Zhang, Zheng-Fu He, Lu-Lu Zhang et al.· Proceedings of the 32nd ACM...· 0 citations