When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but...
SciR is the first multi-paradigm scientific-reasoning benchmark with parametric control on both extraction and inference difficulty, and is believed to be the first multi-paradigm scientific-reasoning benchmark with parametric control on both extraction and inference difficulty.
P. Beckmann, Marco Valentino, André Freitas· arXiv.org· 0 citations
TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map, is demonstrated, through mechanistic analysis and causal interventions, that the model represents intersections and streets, tracks its position, and uses a goal compass to na...
P. Beckmann, Matthieu Queloz, André Freitas· 0 citations
This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which is called feature recall, and defines it, shows it applies across architectures, and contrasts it with the established paradigm of feature combination.
This work argues for the virtual instance view on the grounds that attention streams sustain quasi-psychological connections across token-time, and presents the persona literature, organised around three hypotheses about the internal structure underlying personas in LLMs, and shows that the two persona-based views are...
P. Beckmann, Patrick Butlin· arXiv.org· 8 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.