One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literatu...
Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel et al.· 0 citations
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discov...
Leon Bergen, Usha Bhalla, Andrew Lee et al.· 0 citations
This work introduces axis-aligned feature accentuation, which converts each model’s fitted encoding axis into graded stimulus perturbations that are predicted to parametrically control neural firing within and beyond the natural-image range.
Jacob S. Prince, Binxu Wang, Thomas Fel et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.